
Platform Architect, AI/ML Infrastructure, GCP-focused
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Brazil, +4 more countries.
• Develop and manage infrastructure for model and inference serving for both real-time and batch inference across various tenants.
• Oversee latency, throughput, autoscaling, and reliability measures.
• Take ownership of the ML deployment lifecycle, which includes model registry, versioning, promotion workflows, canary/shadow/A-B rollouts, and rollback processes.
• Manage agentic and LLM workloads in production settings, including oversight of providers, gateways, quotas, throttling, guardrails, prompt/version management, and graceful degradation.
• Create reproducible training, evaluation, and deployment pipelines as code, ensuring lineage and reproducibility.
• Expand Terraform infrastructure-as-code practices for ML systems and multi-project setups.
• Implement GitOps for ML workloads utilizing ArgoCD for configuration and promotion workflows.
• Execute ML and AI workloads on a multi-tenant GKE, focusing on GPU scheduling, workload placement, tenant isolation, and cost-aware capacity management.
• Ensure ML reliability and observability, including inference SLOs, model/data drift detection, regression monitoring, alert quality, on-call ergonomics, and runbooks.
• Promote ML cost efficiency through optimized accelerator sizing, committed-use and Spot VM capacity management, and tenant/workload cost attribution.
• Leverage agentic coding tools to create environments, generate/review IaC and pipeline code, and enhance automation.
• Proactively identify platform issues and contribute to the evolution of the platform.
• 5+ years of experience in platform engineering, SRE, MLOps, or infrastructure, with significant time spent operating production systems at scale.
• Practical experience in deploying and managing ML or AI workloads in a production environment.
• Strong foundation in SRE/DevOps, including ownership of production reliability, SLOs, post-mortems, and measurable enhancements.
• Extensive knowledge of Terraform, including complex state management, reusable modules, multi-project configurations, and CI-driven plan/apply workflows.
• Robust GitOps experience, particularly with ArgoCD or Flux in production settings.
• In-depth understanding of Kubernetes, covering production cluster operations, failure modes, and control-plane intricacies.
• Prior experience with production GKE is highly preferred.
• Strong background in GCP: VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design.
• Hands-on production experience with BigQuery, specifically in partitioning, clustering, query cost/performance tuning, and dataset-level IAM.
• Familiarity with Dataflow, Pub/Sub, or Dataproc.
• Experience in building and managing CI/CD pipelines, with an understanding of the differences in ML pipelines.
• Senior-level approach focused on automation.
• Active engagement with agentic coding tools.
• Excellent written and verbal communication skills.
• Bachelor's degree or equivalent experience in IT or Computer Science as indicated in application questions.
• Availability to work in the EST time zone.
• Preferred experience in GPU/accelerator scheduling, LLM inference at scale, ML orchestration, model registries, drift monitoring, FinOps, data infrastructure, multi-tenant infrastructure, and scaling from startup to enterprise.
• Compensation in USD.
• Remote work opportunity available in LATAM.
• Working hours aligned with the EST time zone.
RR Donnelley
plotdesk
CmdScale GmbH
Colsubsidio
Get handpicked remote jobs straight to your inbox weekly.