
Technical Support Engineer – Inference
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in United States.
• Interact directly with customers to address intricate technical issues related to GPU clusters, inference, and fine-tuning services.
• Serve as a customer-facing Site Reliability Engineer (SRE) to ensure the health, stability, and performance of customer inference endpoints on Kubernetes.
• Develop expertise in the product and act as the final technical defense before escalation to Engineering and Product teams.
• Assist in hardware and platform migrations by assessing system health and traffic management.
• Monitor dashboards, identify anomalies, and escalate issues with data-supported analysis.
• Manage customer communications during incidents and service degradations.
• Convert technical findings into clear, evidence-supported updates for customers.
• Implement infrastructure changes via pull requests and infrastructure-as-code for endpoint configuration, model deployment, capacity scaling, and cluster setup.
• Report engine-level bugs with detailed logs and reproduction steps.
• Collaborate with Engineering, Research, Product, Sales, Support, and senior leadership to address customer issues.
• Recognize support-case trends and contribute to shaping Together AI’s product roadmap.
• Maintain comprehensive documentation that includes system configurations, procedures, troubleshooting guides, and FAQs.
• Provide support coverage during holidays, nights, and weekends as necessary.
• Over 6 years of experience in a customer-facing technical position, such as SRE, DevOps, or infrastructure engineering, with at least one year supporting an AI service.
• Experience as an SRE or DevOps engineer with Kubernetes.
• Strong understanding of AI, ML, GPU technologies, and high-performance computing environments.
• Production-level experience with Kubernetes, SLURM, Ansible, high-performance network fabrics, NFS storage, and container infrastructure.
• Familiarity with Vast and Weka storage solutions in HPC settings.
• Capability to diagnose intricate network-layer problems and interpret traces.
• Proficiency in Python, TypeScript, and/or JavaScript, along with experience using curl and Postman for testing/debugging.
• Expertise with Prometheus and Grafana at scale.
• Deep understanding of REST API debugging and HTTP semantics.
• Experience with LLM inference frameworks, LoRA fine-tuning, and common training failure scenarios.
• Familiarity with Infrastructure as Code and Git-based workflows.
• Background in GPU cluster management.
• Experience with AWS, GCP, and/or Azure.
• Knowledge of compute-cluster installation, configuration, administration, troubleshooting, and security.
• Strong problem-solving and troubleshooting skills for complex technical issues.
• Ability to collaborate cross-functionally with Sales, Engineering, Support, Product, and Research teams.
• Strong sense of ownership and eagerness to learn.
• Excellent communication and interpersonal skills, including the ability to explain complex technical concepts to non-technical stakeholders.
• Capability to manage multiple projects, switch contexts, and prioritize tasks effectively.
• Availability to work during US daytime hours, including Saturdays and Sundays plus two additional weekdays.
• Capability to work a four-day, ten-hour shift with additional weekend on-call responsibilities.
• Flexibility to provide support during holidays, nights, and weekends.
• Startup equity.
• Health insurance.
• Additional benefits.
• Flexibility regarding remote work.
• Competitive compensation.
Immersive Gamebox
Medlogix
alt.bank
Get handpicked remote jobs straight to your inbox weekly.