
MLOps Engineer, LLM Systems, Serving, GPU Kernels, Profiling
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States, +2 more countries.
• Create complex, domain-specific tasks related to GPU kernels, performance profiling, debugging, and inference serving, while producing precise and well-organized solutions.
• Support research and engineering teams in bridging knowledge gaps and enhancing AI model performance across ML systems, training infrastructure, and framework-related subjects.
• Assess MLOps and ML systems tasks and solutions, offering clear, written technical feedback.
• Formulate guidelines and comprehensive evaluation frameworks that address kernel-level optimization, interpretation of profiler outputs, reasoning in distributed systems, and the trade-offs of serving throughput and latency.
• Partner with subject matter experts to ensure training data remains accurate and consistent.
• Contribute to the training and evaluation of AI models by developing and reviewing MLOps and ML systems tasks and solutions for cutting-edge AI training data.
• Over 2 years of practical, hands-on experience in ML systems, ML infrastructure, model serving, or performance engineering related to GPUs and accelerators.
• Experience in at least one of the following areas: writing or optimizing custom GPU kernels (CUDA, Triton, Pallas); performance profiling and trace analysis (Kineto, torch.profiler, Nsight, XLA or JAX profiler); debugging distributed or accelerator-dependent workloads; or serving large language models at scale (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching).
• Proven production experience with JAX and/or PyTorch.
• Familiarity with contemporary accelerators such as A100, H100, B200, or TPU.
• Ability to analyze throughput, latency, and memory trade-offs effectively.
• Evidence of career advancement.
• Commitment to consistently engage for a minimum of 40 hours per week on weekdays.
• Excellent written communication skills, with the ability to articulate complex technical concepts clearly.
• [Benefit details can be added here]
• [Additional benefits can be included here]
SSC HR Solutions
Teamficient
Omilia - Conversational Intelligence
Slate Auto
Get handpicked remote jobs straight to your inbox weekly.