
Senior Platform Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in Canada.
β’ You will ensure the reliability, security, and operability of Element451's platform while developing the delivery systems that facilitate scaling without the need for constant firefighting. This role is a hands-on senior individual contributor position with a broad scope, encompassing core aspects of reliability and operations, as well as CI/CD, delivery, security, infrastructure, and data reliability.
β’ You are responsible for maintaining the operational health of the platform, ensuring it is available, fast, observable, and secure in production. You will also create the automation that makes these attributes sustainable rather than reliant on extraordinary efforts.
β’ Engage in on-call duties, lead incident responses, conduct blameless post-incident reviews, and work towards identifying root causes rather than merely addressing symptoms.
β’ Approach operational toil as an engineering challenge to overcome β consistently automate remediation, enhance alert quality, and reduce MTTD and MTTR instead of accepting manual tasks.
β’ Collaborate with the Director of Platform Engineering to own and enhance the CI/CD and delivery platform β develop pipelines, deployment automation, environment management, and release tools β ensuring that shipping is routine, secure, and low-drama, with features like progressive delivery, automated rollback, and production validation gates as standard.
β’ Serve as the platform's hands-on security operator, focusing on IAM and least-privilege practices, secrets management, threat detection and response (WAF, GuardDuty), and vulnerability triage and remediation in accordance with SLA.
β’ A minimum of 7 years of experience in site reliability, operations, delivery, infrastructure, or platform engineering, demonstrating a history of hands-on delivery β including genuine startup experience where you managed a broad and dynamic scope without a large supporting team.
β’ A robust foundation in SRE principles, including SLI/SLO design, ownership of observability stacks, and incident response in production environments impacting real customers.
β’ Strong expertise in CI/CD and delivery engineering β building and managing pipelines (GitHub Actions), Docker/ECR workflows, and ECS deployment automation, with progressive delivery and automated rollback capabilities β sufficient to co-own the delivery platform, not merely use it.
β’ In-depth knowledge of security operations that you can manage independently, without a dedicated security team β expertise in IAM governance, least-privilege principles, secrets management, network security, threat detection (WAF, GuardDuty), and vulnerability triage and remediation. Your ability to establish operational standards, rather than just adhere to them, is crucial.
β’ Extensive, up-to-date expertise in AWS services β ECS/Fargate, Lambda, SQS/SNS, EventBridge, S3/CloudFront, VPC networking, IAM, and Secrets Manager β along with a strong understanding of Terraform and infrastructure-as-code practices across multi-environment systems.
β’ Operational experience with MongoDB Atlas or a similar managed database platform, including backup and recovery, performance optimization, and large-scale monitoring.
β’ Familiarity with compliance operations β SOC 2 Type II and FERPA β and the ability to produce audit evidence (we utilize Vanta) as a natural byproduct of effective engineering rather than as a separate task.
β’ Ability to operate as a high-output individual contributor with extensive domain ownership and a long-term execution focus. Current knowledge of AI-assisted operations, such as intelligent alerting, anomaly detection, or AI-augmented incident response, is advantageous.
β’ Competitive compensation and comprehensive benefits β a salary aligned with the seniority of the role, along with full medical and dental coverage for you and your family.
β’ Truly remote work environment β we are designed to be remote-first, not just remote-tolerant.
β’ Time to recharge β flexible PTO, paid company holidays, milestone rewards for your tenure, and your birthday off.
β’ Meaningful work β every release contributes to helping students find their path to college and succeed thereafter. Your skills have a tangible impact on people's lives.
Arctiq
Cisco
Prove
Hello Heart
Get handpicked remote jobs straight to your inbox weekly.