Remotery

Senior AI Infrastructure Software Engineer – DGX Cloud

Posted Aug 7

This is a fully remote position, open to applicants in California, +1 more state.

πŸ“‹ Description

β€’ Create platforms and tools for extensive AI, LLM, and GenAI infrastructure.

β€’ Design and enhance tools to boost the efficiency and resilience of AI/ML workloads.

β€’ Diagnose, analyze, and resolve failures from the application level to the hardware level.

β€’ Improve the infrastructure and products that support NVIDIA's AI platforms.

β€’ Collaborate in designing and implementing APIs for integration with NVIDIA's resiliency stacks on the platform.

β€’ Establish meaningful and actionable reliability metrics to monitor and enhance system and service reliability.

β€’ Construct and expand distributed systems and AI infrastructure services.

β€’ Facilitate large-scale AI training, inference, fine-tuning, and Agentic AI in a production environment.


⛳️ Requirements

β€’ A minimum of 8+ years of experience in developing software infrastructure for large-scale AI systems.

β€’ A Bachelor's degree or higher in Computer Science or a related technical field, or equivalent experience.

β€’ Strong debugging abilities and experience in analyzing and addressing AI applications from both application and hardware levels.

β€’ A proven history of building and scaling large distributed systems.

β€’ Experience in AI training, inference, and data infrastructure services.

β€’ Familiarity with Kubernetes.

β€’ Experience managing large-scale observability platforms for monitoring and logging, such as ELK, Prometheus, and Loki.

β€’ Proficiency in Golang, Python, C/C++, and scripting languages.

β€’ Excellent communication and collaboration skills.

β€’ Experience working with large-scale AI clusters and cloud-native infrastructure.

β€’ A strong understanding of NVIDIA GPUs and networking technologies including RDMA, IB, and NCCL.

β€’ Knowledge of deep learning frameworks such as PyTorch, TensorFlow, JAX, Dynamo, and Ray.

β€’ Experience performing root-cause analysis of failures at the datacenter scale.

β€’ A solid background in software design and development.


🏝️ Benefits

β€’ Equity.

β€’ Comprehensive benefits.

β€’ Support and mentorship.

β€’ A culture of blameless postmortems, iterative improvement, and risk-taking.

β€’ An inclusive work environment.

β€’ An equal opportunity employer.

People also viewed

Cloudera16 hours ago

Staff Software Engineer, Flink/Streaming Analytics

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job
Stellar Cyber16 hours ago

Senior / Staff Software Engineer – Parser Team

US flagUnited States OnlyFull-timeFull-stack Engineer$150k – $200k/year
ApplyView job
Pragmatike17 hours ago

Senior Founding Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike17 hours ago

Staff Product Engineer

US flagCalifornia, +3 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike17 hours ago

Founding Product Engineer – YCombinator Experience

US flagCalifornia OnlyFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job
Pragmatike17 hours ago

Lead Product Engineer

US flagCalifornia, +2 more statesFull-timeFull-stack Engineer$200k – $350k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers