
Senior Software Engineer, AI Infrastructure
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design and develop the infrastructure for serving LLMs on Kubernetes, encompassing deployment, GPU scheduling, scaling, and management of the model lifecycle.
• Prepare the platform for enterprise settings utilizing Helm for installations, upgrades, and support for restricted or offline networks.
• Seamlessly integrate the serving layer with the platform's API gateway, identity services, and metering functionalities.
• Establish observability for GPU inference operations in a production environment, including metrics related to serving and GPU telemetry.
• Contribute to a diverse multi-service codebase.
• Influence engineering direction through the preparation of design documents and conducting code reviews.
• Become a member of an early-stage, small senior team with extensive ownership over the model-serving layer and its transition to production.
• A minimum of 5 years of experience in software engineering, focusing on infrastructure, platforms, or distributed systems.
• Extensive hands-on experience with Kubernetes, particularly in building and managing production workloads and Helm charts.
• Familiarity with GPU workloads or LLM inference, or a strong background in related systems along with a proven ability to learn quickly.
• Proficient in Go programming.
• Strong skills in CI/CD and infrastructure-as-code practices.
• Comfortable using AI-assisted development tools like Claude Code and OpenAI Codex within daily engineering tasks.
• Ability to work with a high degree of autonomy on a small, remote-first team that values written communication.
• Nice to have: experience in enhancing inference performance, including quantization, batching, or caching techniques.
• Nice to have: background in distributed serving frameworks.
• Nice to have: experience with enterprise-level deployments, including air-gapped installations, SSO/OIDC, and supply-chain security measures.
• Nice to have: UI development experience, particularly with React/TypeScript.
• Nice to have: contributions to open-source projects within the Kubernetes or ML-infrastructure communities.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company-sponsored outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
• Collaborate with enthusiastic colleagues and Fortune 500 as well as Global 2000 clients.
• Engage in cutting-edge, open-source innovation.
• Thrive in a high-energy atmosphere that promotes openness, collaboration, risk-taking, and ongoing growth.
Alzheimer's Association®
Capital One
Capital One
Goods & Services
Get handpicked remote jobs straight to your inbox weekly.