Experience architecting and operating Kubernetes platforms on bare metal at scale, including lifecycle management, cluster provisioning, and day-2 operations (patching, upgrades, failure recovery) in air-gapped or high-security environments.
Experience with SpectroCloud PaletteAI or similar bare-metal Kubernetes lifecycle management platforms is strongly preferred.
Hands-on experience with NVIDIA AI Enterprise, Run:AI, or equivalent GPU orchestration and scheduling platforms for large-scale AI/ML workloads, including multi-tenant resource management, gang scheduling for distributed training, and GPU fleet health monitoring.
Demonstrated experience integrating multiple third-party software products into a unified platform stack, including managing cross-vendor compatibility, version dependencies, and developing cohesive operational procedures across components from OS through application layer.
Familiarity with high-performance networking in GPU cluster environments and network automation for multi-tenant isolation, including zero-trust network architectures and micro-segmentation approaches for securing multi-tenant AI infrastructure.
Experience developing and executing incremental deployment strategies for complex platforms, including automated provisioning, infrastructure-as-code, and validation testing at progressively larger scale in environments where no prior playbook exists.