
Senior Platform Architect
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in Brazil, +4 more countries.
β’ Construct and manage scalable backend and AI infrastructure for both real-time and batch processing tasks.
β’ Create and uphold deployment workflows that include version control, staged rollouts, automated releases, monitoring, and safe rollback strategies.
β’ Develop and oversee production-level LLM and agentic systems, integrating model providers, APIs, gateways, tools, and external services.
β’ Create reusable services, APIs, automation, and data pipelines for AI-driven products and internal platform functionalities.
β’ Enhance infrastructure-as-code utilizing Terraform and reusable provisioning patterns.
β’ Sustain GitOps deployment workflows using tools such as ArgoCD.
β’ Manage distributed workloads on Kubernetes (GKE), overseeing scaling, workload placement, tenant isolation, service reliability, and infrastructure capacity.
β’ Enhance observability and reliability through metrics, logging, tracing, SLOs, alerting, incident response, and operational tools.
β’ Identify enhancements for performance and infrastructure costs across cloud services, compute, APIs, and AI workloads.
β’ Utilize agentic coding and AI development tools to expedite engineering tasks, review code, debug, and automate repetitive workflows.
β’ Influence the evolution of the platform as scaling, architectural domains, and ML/AI requirements advance.
β’ Over 5 years of experience in platform engineering, SRE, or infrastructure, including managing production systems at scale.
β’ Solid foundation in SRE/DevOps; experience in owning reliability, defining and measuring SLOs, conducting post-mortems, and driving measurable enhancements.
β’ Extensive expertise in Terraform, including complex state management, reusable modules, multi-project configurations, and CI-driven plan/apply workflows.
β’ Strong background in GitOps with ArgoCD or Flux in a production environment.
β’ Comprehensive knowledge of Kubernetes and production cluster operations, including troubleshooting at the control-plane level.
β’ Strong background in cloud infrastructure across AWS, Azure, or GCP, including networking, compute, IAM, storage, and designs for multi-account or multi-project setups.
β’ Practical experience with CI/CD pipelines using GitHub Actions, Cloud Build, GitLab CI, or similar tools.
β’ Senior-level automation-first mindset.
β’ Frequent use of agentic coding tools.
β’ Excellent communication skills for articulating operational decisions, technical trade-offs, and incident summaries.
β’ Bachelor's Degree or equivalent experience as indicated in the application questions.
β’ Degree in IT, Computer Science, or equivalent experience as indicated in the application questions.
β’ Working hours aligned with the EST time zone.
β’ Nice-to-have: Experience with MLOps, GCP, BigQuery, Dataflow, Pub/Sub, Dataproc, GPU scheduling, LLM inference, ML orchestration, model registries, feature stores, drift monitoring, FinOps, data infrastructure, multi-tenant infrastructure, and scaling in startups.
β’ Compensation in USD.
β’ Remote work opportunities available in LATAM.
β’ Working hours aligned with the EST time zone.
RR Donnelley
plotdesk
CmdScale GmbH
Colsubsidio
Get handpicked remote jobs straight to your inbox weekly.