
Senior Solutions Engineer, AI/HPC Networking
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in United States.
• Develop resilient AI/HPC infrastructure for both new and existing clients.
• Engage directly with platforms based on NVIDIA and AMD technology.
• Address operational and reliability concerns of large-scale AI clusters, focusing on performance at scale, training stability, real-time monitoring, logging, and alerting.
• Manage Linux systems, ranging from GPU-enabled servers to general-purpose computing environments.
• Create and strategize rack layouts and network configurations tailored to client specifications.
• Design and assess automation scripts for network operations, and set up server and switch fabrics.
• Execute data center upgrades and facilitate the seamless deployment of DriveNets solutions.
• Install and configure DriveNets products to ensure optimal performance and customer satisfaction.
• Sustain live services by evaluating and monitoring availability, latency, and overall system health.
• Enhance the complete service lifecycle from conception and design through deployment, operation, and improvement.
• Provide insights to internal teams by reporting bugs, documenting workarounds, and proposing enhancements.
• Collaborate with sales teams and customers on significant opportunities and deployments.
• Present new products to DriveNets sales and support teams as well as customers.
• Conduct technical training sessions and TOIs for support and sales engineers, partners, and customers.
• Work together on product definition by gathering customer requirements and planning the roadmap.
• BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or another engineering discipline, or equivalent experience.
• Over 3 years of experience in network engineering (system/solution).
• More than 3 years of experience in solution architecture/sales engineering, or equivalent, with a vendor, value-added reseller, or system integrator.
• Technical proficiency in data center or high-end enterprise network design, including BGP, EVPN, VXLAN, QoS, and Multicast.
• Expertise in data center design, encompassing networking, compute, and storage.
• Capability to produce extensive technical documentation for external audiences, such as white papers and technical briefs.
• Aptitude for multitasking effectively in a dynamic environment.
• Ability to collaborate with teams located across different geographical areas.
• Strong written and verbal communication skills.
• Competence in collaborating efficiently with executives and engineering teams.
• Willingness to travel domestically and internationally up to 20% of the time.
• Familiarity with AI-related data center infrastructure and networking technologies, including InfiniBand, RoCEv2, PFC, ECN, accelerated computing, GPUs, NICs, and DPUs.
• Understanding of AI/HPC networking infrastructure solutions and high-speed interconnect technologies.
• Knowledge of scale-up technologies such as NVLink and UALink.
• Familiarity with scale-out Ethernet and Enhanced Ethernet technologies, InfiniBand, backend storage connectivity, and the fundamentals of data center operations.
• Knowledge of monitoring tools like Prometheus, Grafana, and ELK Stack.
• Acquaintance with telemetry technologies such as gRPC, gNMI, and OTLP.
• Proven experience with one or more Tier-1 Clouds (AWS, Azure, GCP, or OCI) or emerging NeoClouds.
• Background in cloud-native architectures and software.
• Flexible remote work arrangement.
• Opportunities for travel to customer locations.
• Chance to work with AI/HPC networking infrastructure and NVIDIA/AMD platforms.
• Exposure to large-scale AI infrastructure, hyperscalers, NeoClouds, and enterprise deployments.
• Opportunities for international and cross-functional collaboration.
• Professional involvement in technical training, product launches, and roadmap planning.
Expel
Grafana Labs
Livestock Information Ltd
Salesforce
Get handpicked remote jobs straight to your inbox weekly.