
Senior Software Engineer, AI Infrastructure
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design and develop LLM serving infrastructure within Kubernetes, encompassing deployment, GPU scheduling, scaling, and management of the model lifecycle.
• Prepare the platform for enterprise settings via Helm-based installations, upgrades, and configurations for restricted or offline networks.
• Seamlessly integrate the serving layer with the platform’s API gateway, identity services, and metering functionalities.
• Establish observability for GPU inference operations in production, including serving metrics and GPU telemetry.
• Contribute to a multi-service codebase.
• Assist in establishing engineering direction through the creation of design documents and conducting code reviews.
• Collaborate within a small senior team that has comprehensive ownership of the model-serving layer and its transition to production.
• A minimum of 5 years of experience in software engineering focused on infrastructure, platforms, or distributed systems.
• Extensive hands-on experience with Kubernetes, including building and managing production workloads and Helm charts.
• Familiarity with GPU workloads or LLM inference, or a strong background in related systems with a proven ability to learn quickly.
• Proficient programming skills in Go.
• Strong background in CI/CD practices and infrastructure-as-code methodologies.
• Proficient in using AI-assisted development tools, such as Claude Code and OpenAI Codex, as part of the daily engineering workflow.
• Comfortable working with high autonomy in a small, remote-first team that values written communication.
• Preferred: experience in inference performance improvements such as quantization, batching, or caching.
• Preferred: experience with distributed serving frameworks.
• Preferred: experience in enterprise deployments, including air-gapped installations, SSO/OIDC, and supply-chain security.
• Preferred: UI development experience, particularly with React and TypeScript.
• Preferred: contributions to open-source projects within the Kubernetes or ML-infrastructure communities.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package accompanied by a robust benefits plan.
• Flexible remote work arrangement.
Alzheimer's Association®
Capital One
Capital One
Goods & Services
Get handpicked remote jobs straight to your inbox weekly.