
Platform Architect, AI/ML Infrastructure, GCP-focused
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Brazil, +4 more countries.
• Develop and manage infrastructure for model and inference serving to facilitate real-time and batch inference across various tenants.
• Oversee latency, throughput, autoscaling, and system reliability.
• Take responsibility for the ML deployment lifecycle, which includes model registry, version control, promotion workflows, rollout strategies, and safe rollback procedures.
• Manage agentic and LLM workloads in a production environment, covering inference providers, gateways, quotas, throttling, guardrails, prompt/version management, and graceful degradation.
• Create reproducible, automated training, evaluation, and deployment pipelines as code.
• Expand infrastructure-as-code practices to ML systems utilizing Terraform and a multi-project design approach.
• Implement GitOps for ML workloads and manage ArgoCD configuration and promotion workflows.
• Execute ML and AI workloads on multi-tenant Kubernetes/GKE, handling GPU scheduling, workload distribution, tenant isolation, and capacity management.
• Ensure ML reliability and observability, focusing on inference SLOs, model/data drift detection, regression monitoring, alert quality, on-call ergonomics, and runbooks.
• Promote cost efficiency in ML through accelerator right-sizing, committed-use and Spot VM capacity management, and cost attribution.
• Utilize agentic coding tools to set up environments, generate/review IaC and pipeline code, and enhance automation.
• Proactively identify platform issues and contribute to platform evolution.
• 5+ years of experience in platform engineering, SRE, MLOps, or infrastructure, with significant time spent operating production systems at scale.
• Practical experience in deploying and managing ML or AI workloads in a production setting.
• Strong foundation in SRE/DevOps principles, including accountability for reliability, SLOs, post-mortems, and measurable enhancements.
• Extensive knowledge of Terraform, including managing complex states, reusable modules, multi-project configurations, and CI-driven plan/apply workflows.
• Solid GitOps experience using ArgoCD or Flux in a production environment.
• In-depth understanding of Kubernetes and production cluster operations, including troubleshooting at the control plane level.
• Production experience with GKE is highly preferred.
• Strong familiarity with GCP, encompassing VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design.
• Practical experience with BigQuery in production, including partitioning, clustering, query cost/performance tuning, and dataset-level IAM.
• Knowledge of Dataflow, Pub/Sub, or Dataproc is beneficial.
• Experience in building and managing CI/CD pipelines.
• Understanding the distinctions between ML pipelines and standard application CI/CD processes.
• Senior-level automation-first mindset.
• Active engagement with agentic coding tools.
• Excellent communication skills.
• Experience with GPU/accelerator scheduling and node lifecycle management is a plus.
• Experience with operating LLM inference at scale is a plus.
• Familiarity with ML pipeline and orchestration tools is a plus.
• Knowledge of model registries, feature stores, and experiment tracking is a plus.
• Understanding of model and data drift monitoring and ML-specific observability is a plus.
• Background in FinOps is a plus.
• Familiarity with data infrastructure is a plus.
• Experience with multi-tenant infrastructure is a plus.
• Previous experience scaling startups is a plus.
• Bachelor’s Degree or equivalent experience.
• Degree in IT or Computer Science or equivalent experience.
• Proficiency in conversational English.
• Compensation in USD.
• Flexible remote work arrangement.
• Working hours aligned with the EST time zone.
RR Donnelley
plotdesk
CmdScale GmbH
Colsubsidio
Get handpicked remote jobs straight to your inbox weekly.