
Senior/Staff Platform Engineer
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants in Canada, +2 more countries.
• Design, develop, manage, and enhance production Kubernetes platforms.
• Take ownership of cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.
• Diagnose and resolve issues related to Kubernetes, containers, Linux, networking, and the underlying infrastructure.
• Manage and enhance extensive, highly available infrastructure across cloud, hybrid, virtualized, and/or bare-metal settings.
• Create and maintain production tools and automation utilizing Go, Python, or Java.
• Develop internal services, APIs, integrations, and operational tools.
• Automate repetitive operational tasks to minimize manual intervention on the platform.
• Ensure the reliability and operational health of essential production infrastructure.
• Lead or significantly participate in incident responses and root-cause analysis.
• Define and enhance SLOs, SLIs, alerting, and operational workflows.
• Utilize logs, metrics, traces, profiling tools, and system-level diagnostics.
• Propel advancements in availability, performance, capacity, resilience, and operational readiness.
• Contribute to disaster recovery planning, testing, and ongoing improvements.
• Construct and maintain infrastructure as code using Terraform and associated automation technologies.
• Develop and enhance CI/CD and deployment processes.
• Engage in production cloud or infrastructure migrations, which includes dependency analysis, networking, cutover, rollback, and validation.
• Create and maintain monitoring, metrics, dashboards, alerting, logging, and distributed tracing.
• Collaborate directly with customers and internal engineering teams to troubleshoot issues and implement technical solutions.
• Communicate architectural designs, technical decisions, risks, trade-offs, and progress to technical stakeholders.
• Manage complex infrastructure projects from problem identification to production operation.
• Participate in architecture discussions, RFCs, design reviews, and technical direction setting.
• Mentor engineers and enhance engineering and operational practices.
• Operate autonomously in uncertain situations and take ownership when guidance is lacking.
• Over 10 years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related areas; over 10 years is preferred for Staff-level candidates.
• Extensive hands-on experience managing complex production infrastructure and distributed systems.
• Proven experience in building and operating production Kubernetes platforms, not just deploying applications on existing clusters.
• Production programming experience with Go, Python, or Java.
• Strong background in production reliability, incident response, troubleshooting, and operational management.
• Experience in independently leading complex technical initiatives from an ambiguous starting point to production.
• Experience collaborating directly with technical stakeholders or clients and effectively communicating intricate technical subjects.
• A degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
• In-depth understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behavior.
• Strong Linux fundamentals and hands-on troubleshooting of production systems.
• Comprehensive understanding of networking concepts, including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.
• Production experience with at least one leading cloud platform: AWS, GCP, or Alicloud.
• Infrastructure as code experience at scale using Terraform or similar tools.
• Configuration management and automation experience with tools such as Ansible, Puppet, or similar.
• Strong capabilities in production debugging and root-cause analysis across infrastructure and distributed systems.
• Familiarity with observability using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent.
• Experience with CI/CD infrastructure and modern software delivery methodologies.
• Proficient in Docker and container tools as part of the production workflow.
• Understanding of high availability, capacity planning, disaster recovery, and production resilience.
• Exceptional written and verbal technical communication skills.
• Strong analytical, debugging, and problem-solving skills.
• High degree of ownership and capability to work independently.
• Comfortable making technical decisions and driving initiatives in ambiguous situations.
• Capable of effectively communicating with customers, engineers, and technical leadership.
• Possesses strong technical judgment and the ability to articulate trade-offs.
• Able to lead technically and influence others without formal management authority.
• Comfortable working within a distributed, highly technical team.
• Fully remote work arrangement.
• Pacific Hours (8:00 AM – 5:00 PM PST).
• On-call every 4–5 weeks.
Vesta Software Group
SEITENBAU GmbH
M&T Bank
Dimensional Fund Advisors
Get handpicked remote jobs straight to your inbox weekly.