
Principal Platform Engineer
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in United States.
• Design, construct, and manage essential AWS cloud infrastructure, encompassing compute, networking, IAM, and container orchestration.
• Take ownership of organization-wide CI/CD pipelines, artifact management, environment promotion, progressive rollout, and rollback processes.
• Develop and sustain logging, metrics, distributed tracing, alerting, SLIs, and SLOs.
• Promote infrastructure as code methodologies and oversee Kubernetes clusters, custom controllers, and operators.
• Collaborate with security and compliance teams regarding secrets, network policies, and access controls.
• Create self-service developer tools and platform APIs for product and AI engineering teams.
• Manage incident response tools, on-call practices, and standards for reliability engineering.
• Lead or assist in significant platform incident responses.
• Collaborate with engineering leadership and finance on cost visibility, capacity planning, and optimization efforts.
• Produce documentation, runbooks, and maintain clear platform interfaces.
• Engage in code reviews and advocate for engineering best practices.
• Work during regular business hours in the Eastern US time zone.
• Bachelor's degree or higher in Computer Science or a similar coding-focused technical discipline, or equivalent technical experience.
• Over 8 years of experience in platform, infrastructure, or site reliability engineering.
• Proven experience at the Staff, Principal, or Architect level demonstrated through platform architecture or ownership of critical production systems.
• Extensive hands-on experience with CI/CD, including artifact management, environment promotion, and progressive rollout.
• Strong expertise in AWS, particularly in IAM, networking, and container orchestration.
• Experience in production environments using Kubernetes.
• Familiarity with GitHub and GitHub Actions.
• A successful history of building internal developer platforms or tools that have been embraced by engineering teams.
• Proficiency in coding with Python, Go, or TypeScript.
• Capability to operate within a polyglot environment using Python, .NET, and TypeScript.
• Experience with infrastructure as code tools such as Terraform, CloudFormation, or Pulumi.
• Knowledge of modern observability tools, including Prometheus, New Relic, Grafana, Datadog, or OpenTelemetry.
• Ability to lead discussions regarding infrastructure costs and reliability with engineering leadership and finance.
• Strong written communication and documentation abilities.
• Excellent analytical thinking and problem-solving capabilities.
• Exceptional verbal and written communication skills.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance plans.
• Generous paid time off policy including holidays and vacation days.
• Opportunities for professional development and training.
• Flexible working hours and remote work options.
Quantiphi
Agility Technologies Inc
American College of Education
First Due
Get handpicked remote jobs straight to your inbox weekly.