
Platform Engineer – Infrastructure
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in North America.
• Automate the complete server lifecycle from start to finish, encompassing provisioning, operating system installations, configuration, hardware validation, and decommissioning.
• Develop and sustain tools for fleet interaction so that engineers do not have to manually configure machines repeatedly.
• Standardize the software build and deployment processes to machines using CI pipelines, artifacts, deployment tools, and rollback mechanisms.
• Implement security roadmap initiatives, which include SSO-integrated access, jump hosts, secrets management, least-privilege access, and maintaining audit trails.
• Assess fleet costs, including cost per workload, provider comparisons, and decisions regarding hardware refresh, and provide recommendations.
• Enhance infrastructure observability and engage in on-call duties and incident responses.
• Create runbooks and documentation for a remote team.
• Collaborate with the Platform team to support Helius's bare-metal global edge network and high-performance services.
• A minimum of 3 years of experience in infrastructure, DevOps, SRE, or platform engineering.
• Practical experience with bare metal, including provisioning and managing physical servers (PXE, IPMI/BMC, RAID, NICs).
• Familiarity with colocation or dedicated server providers.
• Ability to troubleshoot issues related to hardware versus software.
• Strong understanding of Linux fundamentals, including systemd, networking, storage, and performance tuning.
• Experience in configuration management and infrastructure-as-code (e.g., Ansible, Terraform, Salt, Nix).
• Proficient scripting or programming skills in Python, Go, Bash, or similar languages.
• A security-focused mindset that encompasses access control, blast radius, and secrets management.
• Experience with Rust or a willingness to learn and debug Rust services (preferred).
• Background in high-performance networking or low-latency systems (preferred).
• Familiarity with tools like ClickHouse, Prometheus, or Grafana (preferred).
• Experience at a high-scale B2B infrastructure company (preferred).
• Meaningful equity.
• Generous vacation policy.
• Wellness budgets.
• Support for learning opportunities and travel.
• Flexible, fully remote work environment.
Second Nature
Headway
SYNCREON
Rentokil Pest Control North America
Get handpicked remote jobs straight to your inbox weekly.