
Staff AI Platform Engineer, Infrastructure Services
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Take ownership of, design, secure, and scale the Kong AI Gateway deployment on Konnect Hybrid utilizing KCP/EKS.
• Implement and oversee authentication processes, consumer tiers and budgets, rate limiting, semantic caching, and observability.
• Lead efforts in reliability and incident response, including root-cause analysis, remediation, monitoring, and alerting.
• Architect platform-wide solutions across Jenkins/JPAAS CI/CD, ArgoCD GitOps, Kubernetes, Artifactory/Xray, GitHub Enterprise, and GitHub Actions runners.
• Assess and deploy AI developer tools such as Qodo and LinearB, including recommendations for build-vs-buy decisions.
• Establish technical direction, define architecture and standards, review designs, and provide mentorship to engineers.
• Collaborate with security, DevEx, and product engineering teams to convert needs into platform capabilities.
• Deploy and maintain self-hosted/open-weight model serving infrastructure, focusing on GPU capacity planning, autoscaling, and cost/performance optimization.
• Support LLMOps practices, which include model versioning, evaluation, safe rollout, vector stores, and embedding pipelines.
• Develop observability for token usage, latency, and expenditures across API-based and self-hosted models.
• A minimum of 8 years of experience in platform, infrastructure, or DevOps engineering, managing production systems from end to end.
• Practical experience with API gateway technologies such as Kong, Envoy, or Apigee.
• Strong expertise in Kubernetes and GitOps, including tools like ArgoCD or similar.
• Experience in CI/CD with Jenkins pipeline design and administration, build infrastructure, and management of runner/agent fleets.
• Familiarity with Artifactory, Xray, or comparable artifact and package management systems.
• Experience managing GitHub Enterprise or similar source control platforms.
• Proficient in Terraform and AWS/EKS.
• Background in deploying and operating self-hosted LLM inference stacks such as vLLM, NVIDIA Triton/NIM, TGI, or Ollama.
• Experience with GPU-backed infrastructure, Kubernetes GPU scheduling, and autoscaling practices.
• Understanding of LLMOps practices, including model versioning, evaluation harnesses, and usage/cost observability.
• Proven track record of establishing technical direction, leading cross-team initiatives, and mentoring engineers.
• Excellent and proactive communication skills with both technical and non-technical stakeholders.
• Experience with AI-assisted developer tools at scale is a plus.
• Familiarity with Okta/OIDC and enterprise authentication patterns is preferred.
• Experience with LinearB or similar productivity metrics tools and Qodo or comparable AI code review tools is desired.
• Knowledge of vector databases and RAG pipelines is a plus.
• Exposure to LoRA/QLoRA or analogous fine-tuning pipelines is beneficial.
• Restricted Stock Units (RSUs).
• Employee Stock Purchase Plan (ESPP).
• Flexible time off.
• Paid company holidays and sick leave.
• Gender-neutral parental leave.
• Grandparent leave.
• Medical, dental, and vision insurance coverage.
• 401(k) retirement plan with company matching.
• Life and disability insurance.
• Health and dependent care FSA.
• Optional benefits (hospital, accident, critical illness).
• Employee Assistance Program (EAP).
• ARAG pre-paid legal services.
• Nationwide pet insurance.
• Cancer Care program.
• Global business travel medical insurance.
• Home office allowance.
• Mobile phone reimbursement.
• Wellness coach.
• Wellness/gym reimbursement.
• Fertility coverage.
• Adoption and surrogacy reimbursement.
adconova GmbH
Teleperformance
Trilon Group
Carbon60
Get handpicked remote jobs straight to your inbox weekly.