Senior/Staff Platform Engineer

Posted 12 hours ago

This is a fully remote position, open to applicants in Canada, +2 more countries.

📋 Description

• Design, develop, manage, and enhance production Kubernetes platforms.

• Take ownership of cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.

• Diagnose and resolve issues related to Kubernetes, containers, Linux, networking, and the underlying infrastructure.

• Manage and enhance extensive, highly available infrastructure across cloud, hybrid, virtualized, and/or bare-metal settings.

• Create and maintain production tools and automation utilizing Go, Python, or Java.

• Develop internal services, APIs, integrations, and operational tools.

• Automate repetitive operational tasks to minimize manual intervention on the platform.

• Ensure the reliability and operational health of essential production infrastructure.

• Lead or significantly participate in incident responses and root-cause analysis.

• Define and enhance SLOs, SLIs, alerting, and operational workflows.

• Utilize logs, metrics, traces, profiling tools, and system-level diagnostics.

• Propel advancements in availability, performance, capacity, resilience, and operational readiness.

• Contribute to disaster recovery planning, testing, and ongoing improvements.

• Construct and maintain infrastructure as code using Terraform and associated automation technologies.

• Develop and enhance CI/CD and deployment processes.

• Engage in production cloud or infrastructure migrations, which includes dependency analysis, networking, cutover, rollback, and validation.

• Create and maintain monitoring, metrics, dashboards, alerting, logging, and distributed tracing.

• Collaborate directly with customers and internal engineering teams to troubleshoot issues and implement technical solutions.

• Communicate architectural designs, technical decisions, risks, trade-offs, and progress to technical stakeholders.

• Manage complex infrastructure projects from problem identification to production operation.

• Participate in architecture discussions, RFCs, design reviews, and technical direction setting.

• Mentor engineers and enhance engineering and operational practices.

• Operate autonomously in uncertain situations and take ownership when guidance is lacking.


⛳️ Requirements

• Over 10 years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related areas; over 10 years is preferred for Staff-level candidates.

• Extensive hands-on experience managing complex production infrastructure and distributed systems.

• Proven experience in building and operating production Kubernetes platforms, not just deploying applications on existing clusters.

• Production programming experience with Go, Python, or Java.

• Strong background in production reliability, incident response, troubleshooting, and operational management.

• Experience in independently leading complex technical initiatives from an ambiguous starting point to production.

• Experience collaborating directly with technical stakeholders or clients and effectively communicating intricate technical subjects.

• A degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

• In-depth understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behavior.

• Strong Linux fundamentals and hands-on troubleshooting of production systems.

• Comprehensive understanding of networking concepts, including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.

• Production experience with at least one leading cloud platform: AWS, GCP, or Alicloud.

• Infrastructure as code experience at scale using Terraform or similar tools.

• Configuration management and automation experience with tools such as Ansible, Puppet, or similar.

• Strong capabilities in production debugging and root-cause analysis across infrastructure and distributed systems.

• Familiarity with observability using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent.

• Experience with CI/CD infrastructure and modern software delivery methodologies.

• Proficient in Docker and container tools as part of the production workflow.

• Understanding of high availability, capacity planning, disaster recovery, and production resilience.

• Exceptional written and verbal technical communication skills.

• Strong analytical, debugging, and problem-solving skills.

• High degree of ownership and capability to work independently.

• Comfortable making technical decisions and driving initiatives in ambiguous situations.

• Capable of effectively communicating with customers, engineers, and technical leadership.

• Possesses strong technical judgment and the ability to articulate trade-offs.

• Able to lead technically and influence others without formal management authority.

• Comfortable working within a distributed, highly technical team.


🏝️ Benefits

• Fully remote work arrangement.

• Pacific Hours (8:00 AM – 5:00 PM PST).

• On-call every 4–5 weeks.

People also viewed

Vesta Software Group8 hours ago

Platform Engineer

GB flagUnited Kingdom OnlyFull-timePlatform Engineer
ApplyView job
SEITENBAU GmbH8 hours ago

Platform Reliability Engineer

DE flagGermany OnlyFull-timePlatform Engineer
ApplyView job
M&T Bank9 hours ago

Principal AI Platform Engineer

US flagAlabama OnlyFull-timePlatform Engineer$139.7k – $232.9k/year
ApplyView job
Dimensional Fund Advisors12 hours ago

Senior Database Platform Engineer

US flagNorth Carolina, +1 more stateFull-timePlatform Engineer
ApplyView job
Diakonisches Werk Worms-Alzey13 hours ago

In-House DevOps / Platform Engineer

DE flagGermany OnlyFull-timePlatform Engineer€70k/year
ApplyView job
Testsieger.de14 hours ago

Senior System Engineer – Platform & Security

PT flagPortugal OnlyFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers