Warning: Failed to get DSC: the server could not find the requested resource Initial Kueue managementState: === RUN TestDefaultClusterTrainingRuntimes test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Smoke' --- SKIP: TestDefaultClusterTrainingRuntimes (0.00s) === RUN TestOpenMPICudaClusterTrainingRuntime cluster_training_runtimes_test.go:121: Skip due to issue RHOAIENG-61966 --- SKIP: TestOpenMPICudaClusterTrainingRuntime (0.00s) === RUN TestDefaultTrainingHubRuntimesMatchDefaultClusterRuntimes test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Smoke' --- SKIP: TestDefaultTrainingHubRuntimesMatchDefaultClusterRuntimes (0.00s) === RUN TestRunTrainJobWithDefaultClusterTrainingRuntimes test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier1' --- SKIP: TestRunTrainJobWithDefaultClusterTrainingRuntimes (0.00s) === RUN TestJobSetWorkflow test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier1' --- SKIP: TestJobSetWorkflow (0.00s) === RUN TestFailedJobSetWorkflow test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestFailedJobSetWorkflow (0.00s) === RUN TestKubeflowSdk test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier1' --- SKIP: TestKubeflowSdk (0.00s) === RUN TestKubeflowSdkKueueIntegration test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier1' --- SKIP: TestKubeflowSdkKueueIntegration (0.00s) === RUN TestKubeflowSdkOpenMPICudaKueueIntegration kubeflow_sdk_test.go:40: Skip due to issue RHOAIENG-61966 --- SKIP: TestKubeflowSdkOpenMPICudaKueueIntegration (0.00s) === RUN TestSftTrainingHubSingleNodeSingleGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestSftTrainingHubSingleNodeSingleGPU (0.00s) === RUN TestOsftTrainingHubSingleNodeSingleGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestOsftTrainingHubSingleNodeSingleGPU (0.00s) === RUN TestLoraTrainingHubSingleNodeSingleGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestLoraTrainingHubSingleNodeSingleGPU (0.00s) === RUN TestOsftTrainingHubMultiNodeMultiGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestOsftTrainingHubMultiNodeMultiGPU (0.00s) === RUN TestLoraTrainingHubMultiNodeMultiGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestLoraTrainingHubMultiNodeMultiGPU (0.00s) === RUN TestSftTrainingHubMultiNodeMultiGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestSftTrainingHubMultiNodeMultiGPU (0.00s) === RUN TestRhaiTrainingProgressionCPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestRhaiTrainingProgressionCPU (0.00s) === RUN TestRhaiJitCheckpointingCPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestRhaiJitCheckpointingCPU (0.00s) === RUN TestRhaiFeaturesCPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestRhaiFeaturesCPU (0.00s) === RUN TestRhaiTrainingProgressionCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiTrainingProgressionCuda (0.00s) === RUN TestRhaiJitCheckpointingCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiJitCheckpointingCuda (0.00s) === RUN TestRhaiFeaturesCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiFeaturesCuda (0.00s) === RUN TestRhaiTrainingProgressionRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiTrainingProgressionRocm (0.00s) === RUN TestRhaiJitCheckpointingRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiJitCheckpointingRocm (0.00s) === RUN TestRhaiFeaturesRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiFeaturesRocm (0.00s) === RUN TestRhaiTrainingProgressionMultiGpuCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiTrainingProgressionMultiGpuCuda (0.00s) === RUN TestRhaiJitCheckpointingMultiGpuCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiJitCheckpointingMultiGpuCuda (0.00s) === RUN TestRhaiFeaturesMultiGpuCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiFeaturesMultiGpuCuda (0.00s) === RUN TestRhaiTrainingProgressionMultiGpuRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiTrainingProgressionMultiGpuRocm (0.00s) === RUN TestRhaiJitCheckpointingMultiGpuRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiJitCheckpointingMultiGpuRocm (0.00s) === RUN TestRhaiFeaturesMultiGpuRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestRhaiFeaturesMultiGpuRocm (0.00s) === RUN TestTrainingFailureScenarios test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestTrainingFailureScenarios (0.00s) === RUN TestTorchrunTrainingFailure test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestTorchrunTrainingFailure (0.00s) === RUN TestRhaiS3CheckpointingCPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestRhaiS3CheckpointingCPU (0.00s) === RUN TestRhaiS3FsdpFullStateCheckpointingCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3FsdpFullStateCheckpointingCuda (0.00s) === RUN TestRhaiS3FsdpFullStateCheckpointingMultiProcessCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3FsdpFullStateCheckpointingMultiProcessCuda (0.00s) === RUN TestRhaiS3FsdpSharedStateCheckpointingCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3FsdpSharedStateCheckpointingCuda (0.00s) === RUN TestRhaiS3FsdpSharedStateCheckpointingMultiGpuCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3FsdpSharedStateCheckpointingMultiGpuCuda (0.00s) === RUN TestRhaiS3DeepspeedStage0CheckpointingCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3DeepspeedStage0CheckpointingCuda (0.00s) === RUN TestRhaiS3DeepspeedStage0CheckpointingMultiGpuCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestRhaiS3DeepspeedStage0CheckpointingMultiGpuCuda (0.00s) === RUN TestPyTorchDDPMultiNodeMultiCPUWithTorchCuda28 test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestPyTorchDDPMultiNodeMultiCPUWithTorchCuda28 (0.00s) === RUN TestPyTorchDDPSingleNodeSingleGPUWithTorchCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestPyTorchDDPSingleNodeSingleGPUWithTorchCuda (0.00s) === RUN TestPyTorchDDPSingleNodeMultiGPUWithTorchCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestPyTorchDDPSingleNodeMultiGPUWithTorchCuda (0.00s) === RUN TestPyTorchDDPMultiNodeSingleGPUWithTorchCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestPyTorchDDPMultiNodeSingleGPUWithTorchCuda (0.00s) === RUN TestPyTorchDDPMultiNodeMultiGPUWithTorchCuda test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestPyTorchDDPMultiNodeMultiGPUWithTorchCuda (0.00s) === RUN TestPyTorchDDPSingleNodeSingleGPUWithTorchRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestPyTorchDDPSingleNodeSingleGPUWithTorchRocm (0.00s) === RUN TestPyTorchDDPSingleNodeMultiGPUWithTorchRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestPyTorchDDPSingleNodeMultiGPUWithTorchRocm (0.00s) === RUN TestPyTorchDDPMultiNodeSingleGPUWithTorchRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestPyTorchDDPMultiNodeSingleGPUWithTorchRocm (0.00s) === RUN TestPyTorchDDPMultiNodeMultiGPUWithTorchRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestPyTorchDDPMultiNodeMultiGPUWithTorchRocm (0.00s) === RUN TestKueueWorkloadPreemptionSuspendsTrainJob test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier1' --- SKIP: TestKueueWorkloadPreemptionSuspendsTrainJob (0.00s) === RUN TestKueueWorkloadInadmissibleWithNonExistentLocalQueue test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Tier2' --- SKIP: TestKueueWorkloadInadmissibleWithNonExistentLocalQueue (0.00s) === RUN TestSetupUpgradeTrainJob trainer_kueue_upgrade_training_test.go:83: Skip due to issue RHOAIENG-48867 --- SKIP: TestSetupUpgradeTrainJob (0.00s) === RUN TestRunUpgradeTrainJob trainer_kueue_upgrade_training_test.go:158: Skip due to issue RHOAIENG-48867 --- SKIP: TestRunUpgradeTrainJob (0.00s) === RUN TestSetupSpecificRuntimeUpgradeTrainJob trainer_kueue_upgrade_training_test.go:205: Skip due to issue RHOAIENG-48867 --- SKIP: TestSetupSpecificRuntimeUpgradeTrainJob (0.00s) === RUN TestRunSpecificRuntimeUpgradeTrainJob trainer_kueue_upgrade_training_test.go:295: Skip due to issue RHOAIENG-48867 --- SKIP: TestRunSpecificRuntimeUpgradeTrainJob (0.00s) === RUN TestSetupCustomRuntimeUpgradeTrainJob test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Pre-Upgrade' --- SKIP: TestSetupCustomRuntimeUpgradeTrainJob (0.00s) === RUN TestRunCustomRuntimeUpgradeTrainJob kueue_operator.go:109: SetupKueue: Setting kueue to Unmanaged managementState in DataScienceCluster... kueue_operator.go:111: Should be able to set DSC kueue to Unmanaged Unexpected error: <*errors.StatusError | 0x209fde7b54a0>: the server could not find the requested resource { ErrStatus: { TypeMeta: {Kind: "", APIVersion: ""}, ListMeta: { SelfLink: "", ResourceVersion: "", Continue: "", RemainingItemCount: nil, }, Status: "Failure", Message: "the server could not find the requested resource", Reason: "NotFound", Details: { Name: "", Group: "", Kind: "", UID: "", Causes: [ { Type: "UnexpectedServerResponse", Message: "404 page not found", Field: "", }, ], RetryAfterSeconds: 0, }, Code: 404, }, } occurred kueue_operator.go:99: Collecting Kueue diagnostics... kueue_operator.go:99: Failed to get Kueue CR for diagnostics: the server could not find the requested resource (get kueues.kueue.openshift.io cluster) kueue_operator.go:99: Failed to get DSC for diagnostics: the server could not find the requested resource kueue_operator.go:99: No Kueue operator pods found --- FAIL: TestRunCustomRuntimeUpgradeTrainJob (0.02s) === RUN TestMultiNodeOpenMPITrainJob trainer_mpi_test.go:35: Skip due to issue RHOAIENG-61966 --- SKIP: TestMultiNodeOpenMPITrainJob (0.00s) === RUN TestOpenMPICudaTrainJobKueueIntegration trainer_openmpi_kueue_integration_test.go:36: Skip due to issue RHOAIENG-61966 --- SKIP: TestOpenMPICudaTrainJobKueueIntegration (0.00s) === RUN TestOpenMPICudaTrainJobKueueWorkloadDeactivateReactivate trainer_openmpi_kueue_integration_test.go:138: Skip due to issue RHOAIENG-61966 --- SKIP: TestOpenMPICudaTrainJobKueueWorkloadDeactivateReactivate (0.00s) === RUN TestSftStockTrlSingleNodeSingleGPU test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-CUDA' --- SKIP: TestSftStockTrlSingleNodeSingleGPU (0.00s) === RUN TestSftStockTrlSingleNodeSingleGPUWithTorchRocm test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'KFTO-ROCm' --- SKIP: TestSftStockTrlSingleNodeSingleGPUWithTorchRocm (0.00s) === RUN TestKubeflowTrainerSmoke test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Smoke' --- SKIP: TestKubeflowTrainerSmoke (0.00s) === RUN TestSetupTrainingRuntime test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Pre-Upgrade' --- SKIP: TestSetupTrainingRuntime (0.00s) === RUN TestVerifyTrainingRuntime trainer_trainingruntime_upgrade_test.go:76: Failed to list TrainingRuntimes Unexpected error: <*errors.StatusError | 0x209fde69e6e0>: the server could not find the requested resource (get trainingruntimes.trainer.kubeflow.org) { ErrStatus: { TypeMeta: {Kind: "", APIVersion: ""}, ListMeta: { SelfLink: "", ResourceVersion: "", Continue: "", RemainingItemCount: nil, }, Status: "Failure", Message: "the server could not find the requested resource (get trainingruntimes.trainer.kubeflow.org)", Reason: "NotFound", Details: { Name: "", Group: "trainer.kubeflow.org", Kind: "trainingruntimes", UID: "", Causes: [ { Type: "UnexpectedServerResponse", Message: "404 page not found", Field: "", }, ], RetryAfterSeconds: 0, }, Code: 404, }, } occurred test.go:152: Creating ephemeral output directory as TEST_OUTPUT_DIR env variable is unset test.go:160: Output directory has been created at: /tmp/TestVerifyTrainingRuntime893502195 --- FAIL: TestVerifyTrainingRuntime (0.02s) === RUN TestSetupSleepTrainJob test_tag.go:37: Test tier 'Post-Upgrade' doesn't match expected tier 'Pre-Upgrade' --- SKIP: TestSetupSleepTrainJob (0.00s) === RUN TestVerifySleepTrainJob trainer_upgrade_sleep_test.go:70: Expected <[]v1.Pod | len:0, cap:0>: nil to have length 1 test.go:152: Creating ephemeral output directory as TEST_OUTPUT_DIR env variable is unset test.go:160: Output directory has been created at: /tmp/TestVerifySleepTrainJob2330281675 --- FAIL: TestVerifySleepTrainJob (0.09s) FAIL TearDown: Setting kueue managementState to Removed in DataScienceCluster... TearDown: Failed to set Kueue to Removed: TearDown: failed to set kueue to Removed: the server could not find the requested resource ok github.com/opendatahub-io/distributed-workloads/tests/trainer 0.185s