
Senior Platform Engineer
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Canada.
• Transform existing development-only Dockerfiles into production-ready, multi-stage container builds featuring health checks and graceful shutdown management (including draining in-flight background tasks prior to termination)
• Establish and manage a container registry, including policies for access, conventions for image tagging, and automation for retention/cleanup
• Enhance CI/CD pipelines to incorporate container build-and-push steps, transitioning deployment targets from traditional package-based methods to container-based App Service deployments
• Create Infrastructure-as-Code modules from the ground up, encompassing compute, database, storage, secrets, and monitoring resources
• Develop provisioning automation for identity/app-registration setup, database initialization, and environment bootstrapping
• Implement liveness and readiness health checks that confirm connectivity to downstream dependencies (database, queue, background job processing)
• Construct a telemetry pipeline that sends sanitized, PII-free operational metrics to a centralized monitoring system, with adjustable scope and destination
• Design and sustain a tracking/registry system for infrastructure deployments, documenting version, status, and configuration state across multiple environments
• Broaden deployment pipelines to facilitate automated, tag-based promotion and staged/canary rollout strategies, ensuring per-environment failure isolation and rollback
• Develop a controlled process for deploying critical patches outside the standard release schedule, including audit logging, notification workflows, and approval checkpoints
• Collaborate with engineering to review core subsystems (background job processing, database partitioning/sharding, secrets access patterns, external integrations, feature-flag and analytics tools, licensing/validation logic, notifications, scheduled tasks) for portability across deployment environments
• Optimize databases (Azure SQL, Cosmos DB) for enhanced performance and scalability, incorporating sharding and partitioning strategies
• Ensure efficient horizontal scaling strategies to accommodate SaaS growth
• Implement caching solutions (Redis, CDN) and performance tuning methodologies
• Lead incident response and root cause analysis (RCA) initiatives, minimizing mean time to recovery (MTTR)
• Participate in the on-call rotation as a primary responder to production incidents and emergency situations, delivering prompt triage and resolution outside of regular business hours
• Collaborate closely with Engineering, Security, and Product teams to align platform objectives with business goals
• Serve as a technical mentor for junior and mid-level engineers, promoting best practices in DevOps, cloud, and automation
• Advocate for reliability engineering practices (SRE principles, SLAs, SLOs, error budgets)
• 6+ years of experience in DevOps
• A minimum of 4 years of DevOps production experience within a SaaS product-led organization
• Hands-on experience in creating Infrastructure-as-Code from scratch is highly regarded
• Experience transitioning a service from non-containerized to production-quality containerized deployment is highly valued
• Nice-to-have: 3+ years of experience with the .NET platform
• Competitive compensation
• Equity
• Benefits
Quantiphi
Agility Technologies Inc
American College of Education
First Due
Get handpicked remote jobs straight to your inbox weekly.