
Senior AI Platform Engineer – Infrastructure Services
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Assume responsibility for the AI Gateway infrastructure built on Kong AI Gateway, encompassing authentication, routing, rate limiting, and monitoring of AI coding assistant traffic across the organization.
• Design, secure, and scale the deployment of Kong AI Gateway (Konnect Hybrid on KCP/EKS).
• Lead root-cause analysis, remediation efforts, incident response, monitoring, and alerting to ensure gateway reliability.
• Create solutions involving CI/CD, GitOps, Kubernetes deployment tools, artifact management, GitHub Enterprise, and GitHub Actions runner infrastructure.
• Assess and implement AI developer tools, including structured pilot programs and adoption initiatives.
• Establish AI infrastructure architecture and standards, review design proposals, and mentor engineering teams.
• Collaborate with security, DevEx, and product engineering teams to convert requirements into platform capabilities.
• Deploy and manage self-hosted/open-weight model serving infrastructure, including GPU capacity planning, autoscaling, and cost/performance optimization.
• Develop LLMOps practices that include model versioning, evaluation, safe rollout, vector stores, and embedding pipelines.
• Enhance observability for token usage, latency, and expenditures across both API-based and self-hosted models.
• A minimum of 5 years of experience in platform, infrastructure, or DevOps engineering.
• Proven history of owning production systems from start to finish.
• Practical experience with API gateway technologies such as Kong, Envoy, or Apigee.
• Strong background in Kubernetes and GitOps, including experience with ArgoCD or similar tools.
• Experience managing operations across development, government, and production environments.
• Proficiency in designing and administering Jenkins pipelines.
• Experience in managing infrastructure and runner/agent fleets.
• Familiarity with artifact and package management systems like Artifactory, Xray, or similar tools.
• Experience with GitHub Enterprise administration.
• Working knowledge of Terraform and AWS/EKS.
• Experience deploying and managing self-hosted LLM inference stacks and GPU-backed infrastructure.
• Understanding of Kubernetes GPU scheduling and autoscaling principles.
• Familiarity with LLMOps practices, including model versioning, evaluation harnesses, and usage/cost observability.
• Demonstrated ability to set technical direction, lead cross-team initiatives, and mentor engineers.
• Excellent, proactive communication skills for articulating infrastructure trade-offs to both technical and non-technical stakeholders.
• Experience with AI-assisted developer tools at scale is preferred.
• Familiarity with Okta/OIDC and enterprise authentication patterns is preferred.
• Experience with LinearB or similar engineering productivity metrics tools and Qodo or similar AI code review tools is preferred.
• Experience with vector databases and RAG pipelines is preferred.
• Exposure to LoRA/QLoRA or comparable model fine-tuning processes is preferred.
• Restricted Stock Units (RSUs)
• Employee Stock Purchase Plan (ESPP)
• Flexible time off
• Paid company holidays and paid sick leave
• Gender-neutral parental leave
• Grandparent leave
• Medical, dental, and vision insurance
• 401(k) retirement plan with company matching
• Life and disability insurance
• Health and dependent care FSA
• Voluntary benefits (hospital, accident, critical illness)
• Employee Assistance Program (EAP)
• ARAG pre-paid legal services
• Nationwide pet insurance
• Cancer Care program
• Global business travel medical insurance
• Home office allowance
• Mobile phone reimbursement
• Wellness coaching
• Wellness/gym reimbursement
• Fertility coverage
• Adoption and surrogacy reimbursement
adconova GmbH
Teleperformance
Trilon Group
Carbon60
Get handpicked remote jobs straight to your inbox weekly.