Job Description
A leading Kubernetes-native AI infrastructure company is hiring a Senior HPC Networking Engineer to support the design, deployment, and optimization of high-performance networking environments. This role is ideal for experienced network engineers with expertise in InfiniBand technologies, Fortinet solutions, and large-scale HPC or AI infrastructure. The successful candidate will help ensure secure, reliable, and high-performing network operations for data-intensive workloads.
Location: Remote, USA
Employment Type: Full-time
Key Responsibilities
- Design, deploy, and maintain high-performance networking infrastructure with a strong focus on InfiniBand fabrics.
- Troubleshoot complex networking issues across InfiniBand and Ethernet environments to maintain optimal system performance.
- Configure, manage, and optimize InfiniBand switches, subnet managers, HCAs, and fabric components.
- Implement and manage Fortinet security solutions, including FortiGate firewalls, FortiManager, and FortiAnalyzer.
- Monitor network performance, perform capacity planning, and optimize latency and throughput.
- Collaborate with compute, storage, and platform engineering teams to support HPC cluster operations.
- Create and maintain network architecture documentation, configurations, and operational procedures.
- Participate in network upgrades, infrastructure migrations, and on-call support for critical incidents.
Requirements
- Minimum of 5 years of experience in network engineering, preferably within HPC or data center environments.
- Strong hands-on experience with InfiniBand technologies such as Mellanox/NVIDIA.
- Solid understanding of TCP/IP networking, BGP, OSPF, VLANs, QoS, and enterprise network design.
- Experience deploying, configuring, and troubleshooting Fortinet solutions, including FortiGate firewalls, VPNs, and firewall policies.
- Proficiency with network performance monitoring and troubleshooting tools.
- Familiarity with Linux operating systems and scripting languages such as Bash or Python.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Experience supporting large-scale HPC clusters or AI/ML infrastructure.
- Knowledge of RDMA, MPI, and low-latency networking technologies.
- Industry certifications such as FCSS, FCNSP, CCNP, CCIE, or equivalent.
- Experience with automation and Infrastructure as Code tools such as Ansible or Terraform.
Benefits
- Opportunity to work with advanced HPC and AI infrastructure technologies.
- Innovative and collaborative engineering environment.
- Competitive salary and comprehensive employee benefits package.