
Principal Platform Engineer, AI Engineering
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Design and manage a Terraform monorepo across development, QA, staging, and production environments.
• Operate EKS clusters comprehensively, addressing node lifecycle, autoscaling, ingress, workload identity, secrets delivery, and cluster security.
• Develop push-based CI/CD systems utilizing self-hosted GitHub Actions runners with immutable artifacts and strict promotion workflows.
• Maintain shared Helm chart libraries along with per-service charts for backends, frontends, and scheduled jobs.
• Create streamlined paths for launching new services, ensuring logging, metrics, secrets, identity, and pipelines are integrated.
• Implement least-privilege IAM practices, manage secrets, define network boundaries, ensure image provenance, and establish production guardrails.
• Implement cloud cost tagging and allocation strategies, optimize compute resources, and maintain predictable spending as traffic, data, and model inference increase.
• Develop standardized structured logging, metrics, tracing, and cross-service correlation using versioned telemetry contracts.
• Define platform standards for tagging, naming conventions, DNS, versioning, and security; document decisions and review infrastructure and deployment changes.
• Mentor engineers and collaborate with application, data, and AI engineering teams on platform agreements and promotion environments.
• 8+ years of experience building and managing production platform infrastructure (not a strict cutoff; strong candidates with less experience may be considered).
• Practical experience with production Kubernetes, encompassing cluster lifecycle, autoscaling, ingress, workload identity, secrets delivery, and hardening; EKS is preferred.
• Terraform infrastructure-as-code experience at scale, including module design, state organization across environments, and provider upgrades.
• Proven experience in building or significantly reconstructing CI/CD systems, ensuring artifact immutability and a build-once/promote-everywhere delivery model.
• Experience operating self-hosted GitHub Actions runners at scale.
• Hands-on experience with AWS, including IAM, VPC networking and DNS, secrets management, container registries, and managed compute services.
• Extensive Helm experience at scale, including shared chart libraries, templating boundaries, and environmental configuration.
• Familiarity with GitOps or push-based deployment workflows.
• Experience in creating service templates, streamlined paths, self-service tools, and documentation.
• Expertise in implementing observability with structured logging, metrics, and distributed tracing; examples include OpenTelemetry, Prometheus/Grafana, and Datadog.
• Practical security experience in regulated or security-sensitive settings, focusing on least-privilege IAM, secrets management, network isolation, image provenance and scanning, and handling of PHI/PII.
• Experience supporting data workloads on Kubernetes, including tools such as Spark, Kafka, Airflow, or Dagster.
• Experience in cloud cost management, including tagging, allocation, right-sizing, and achieving measurable reductions in spending without compromising reliability.
• Production coding experience and proficiency in at least one backend programming language such as Python, Go, or C#/.NET; shell scripting skills are also required.
• Outstanding communication and collaboration abilities.
• Experience mentoring engineers and establishing extensible infrastructure standards.
• Comfort working within a small, agile team and adapting to multiple roles.
• Comprehensive health and wellness benefits.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
• Competitive salary and performance-based bonuses.
Safari Micro
SysMap Solutions
Flexential
Get handpicked remote jobs straight to your inbox weekly.