Submit Resume

GPU Network Engineer Operational Support

  • Colorado, Englewood

  • 08/25/2026

  • Contract

  • Active

Job Description:

  • Job Summary
    We are seeking a senior GPU Network Engineer to provide Tier 3 operational support for GPU networking environments supporting AI and machine learning workloads. This role will focus on troubleshooting complex issues involving GPU clusters, high-performance networking, RoCE, InfiniBand, and AI infrastructure. The engineer will serve as a technical escalation point and collaborate with network, compute, platform, security, and operations teams to ensure performance, availability, and rapid resolution of production issues. The ideal candidate will have hands-on experience supporting GPU clusters, AI/ML infrastructure, High Performance Computing (HPC) environments, or large-scale low-latency data center networks.

    Key Responsibilities
    • Serve as a Tier 3 escalation resource for GPU network incidents and production issues.
    • Troubleshoot complex networking, performance, and connectivity issues impacting GPU clusters.
    • Support AI, machine learning, and high-performance computing environments.
    • Analyze network performance, congestion, latency, packet loss, and throughput issues.
    • Collaborate with security operations, network, compute, platform, and infrastructure teams during incident response.
    • Validate and optimize GPU network performance for customer workloads.
    • Identify root causes and implement corrective actions for recurring issues.
    • Support production turn-up and go-live activities for new GPU environments.
    • Participate in operational readiness reviews and knowledge transfer sessions.
    • Develop operational runbooks, troubleshooting guides, and support procedures.
    • Provide recommendations for capacity planning, scalability, and operational improvements.

    Required Qualifications
    • 5+ years of data center networking experience.
    • 2+ years of experience supporting GPU, AI/ML, HPC, or large-scale compute infrastructure environments.
    • Experience troubleshooting complex network performance issues in production environments.
    • Strong understanding of data center network architecture and operations.
    • Experience supporting high-bandwidth, low-latency network fabrics.
    • Experience with Arista, NVIDIA, Cisco, or equivalent data center networking technologies.
    • Strong understanding of Layer 2 and Layer 3 networking.
    • Strong understanding of routing and switching.
    • Experience with network monitoring and troubleshooting.
    • Experience with incident management and escalation processes.
    • Ability to work effectively during critical outages and customer-impacting incidents.

    Preferred Qualifications
    • Experience supporting NVIDIA DGX, HGX, SuperPOD, or similar GPU infrastructure.
    • Experience with RDMA, RoCE, InfiniBand, AI network fabrics, or east-west traffic optimization.
    • Knowledge of NVIDIA reference architectures.
    • Experience with Kubernetes-based AI environments.
    • Experience supporting hyperscale, cloud, NeoCloud, or HPC environments.
    • Exposure to optical networking, buffer tuning, congestion management, and performance validation.
    • Familiarity with network telemetry and performance analytics tools.

.

.

.