Job Description
The Senior DevOps Engineer (Storage) will be responsible for deploying, integrating, operating, and optimizing high-performance storage infrastructure for GPU-accelerated computing and AI platforms. The role focuses on the storage layer connecting Kubernetes environments with bare-metal infrastructure, including NFS-based storage, CSI integration, Linux systems, networking, automation, and observability.
Location: Remote, United States
Employment Type: Full-time
Work Arrangement: Remote
Key Responsibilities
- Integrate high-performance NFS-based storage solutions, such as VAST and Dell PowerScale, into Kubernetes clusters using CSI, storage classes, and persistent volumes.
- Optimize NFS data paths, including mount options, nconnect, RDMA, Linux client configurations, and network settings for high-throughput, low-latency GPU and AI workloads.
- Deploy and operate storage services and operators while managing capacity, quotas, snapshots, and storage lifecycles.
- Configure and optimize Linux systems for storage workloads, including drivers, file systems, networking, and kernel parameters.
- Deliver storage integration for k0s-based Kubernetes environments using Cluster API (CAPI) and K0rdent management and child-cluster architectures.
- Operate storage infrastructure in disconnected and air-gapped environments, including local artifact repositories, Harbor connectivity, PKI, and TLS configurations.
- Automate storage provisioning and configuration using infrastructure-as-code tools such as Terraform or OpenTofu and GitOps platforms such as ArgoCD or Flux.
- Develop monitoring, alerting, and observability solutions for storage performance, capacity, reliability, and overall health.
- Investigate and resolve performance, reliability, and scalability issues across the storage and infrastructure stack.
Requirements
- 7+ years of experience in Site Reliability Engineering (SRE) or infrastructure operations.
- 5+ years of experience building and operating distributed production storage systems at scale.
- Hands-on experience with high-performance storage platforms such as VAST, Weka, DDN, or Dell PowerScale.
- Strong understanding of Linux and Kubernetes storage fundamentals, including NFS and CSI.
- Strong Linux storage and networking knowledge, including experience troubleshooting performance at the kernel and NFS-client level.
- Experience with infrastructure-as-code and GitOps practices.
- Ability to independently diagnose complex storage, infrastructure, and reliability issues.
- Strong communication skills and the ability to establish operational standards across teams.
Preferred Qualifications
- Hands-on experience provisioning bare-metal hosts, configuring raw disks and hardware layouts, and managing physical server storage.
- Practical experience with VAST and/or Dell PowerScale.
- Experience with GPUDirect Storage and RDMA/RoCE data paths.
- Experience with the Mirantis K0rdent ecosystem, including K0rdent Enterprise, K0rdent AI, k0s, MKE, and Cluster API.
- Familiarity with additional storage backends such as Ceph and object/S3 storage, including CSI driver operations.
- Experience operating infrastructure in sovereign, high-security, or fully air-gapped environments.
