
OpenStack Infrastructure Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in South Africa.
• Assist in the development, management, and scaling of OpenStack compute and storage infrastructure.
• Collaborate remotely with team members situated in various time zones.
• Add value to the OpenStack team by utilizing a strong understanding of infrastructure systems and expertise in troubleshooting.
• Proven experience in operating open-source infrastructure systems.
• Significant experience is anticipated at the senior level.
• Hands-on experience with OpenStack implementation, operations, or troubleshooting.
• Familiarity with Ceph or comparable distributed storage systems, including aspects such as capacity, replication, failure domains, recovery, latency, and performance.
• Strong skills in Linux systems administration and troubleshooting.
• Practical knowledge of data-center-grade hardware, encompassing enterprise servers, CPUs, memory, storage media, NICs, firmware, and out-of-band management.
• Understanding of how hardware design, power, cooling, rack layout, and component failure impact platform reliability and performance.
• Solid networking fundamentals, with advantageous experience in OVN, OVS, BGP underlays, LACP, Juniper, IPv4, IPv6, and both physical and virtual cloud networks.
• Ability to troubleshoot across various domains including compute, storage, physical and virtual networking, databases, message queues, containers, hypervisors, and guest workloads.
• Experience with virtualization and cloud infrastructure at scale.
• A security-focused approach towards architecture, automation, access control, and operational workflows.
• An SRE mindset aimed at minimizing failure probability, recovery time, operational risks, and data-loss exposure.
• Experience with Infrastructure as Code, configuration management, and automation methodologies.
• Preference for repeatable, version-controlled automation rather than undocumented manual processes.
• Discipline in planning and executing upgrades, migrations, and significant infrastructure changes.
• Strong end-to-end ownership, covering everything from physical infrastructure and network fabric to OpenStack services and customer workloads.
• Capability to evaluate capacity and performance across CPU, memory, storage, IOPS, latency, throughput, packet rates, and control-plane scale.
• Ability to effectively utilize AI-assisted tools while recognizing their limitations.
• Proficient technical documentation skills, including architecture decisions, runbooks, change plans, incident findings, and recovery procedures.
• Capacity to contribute constructively to technical reviews, challenge unsafe assumptions, and respond positively to detailed feedback.
• Competence in working effectively within a distributed, primarily asynchronous team.
• Willingness to participate in an on-call rotation managed through Rootly and to respond to occasional out-of-hours incidents, maintenance, or operational requests as needed.
• Recognizes the distinction between “done” and “97% done” and the potentially significant costs associated with the latter.
• Comprehensive health and wellness programs.
• Opportunities for professional growth and development.
• Flexible work arrangements to promote work-life balance.
Tailscale
Adzuna
Webflow
Coinbase
Get handpicked remote jobs straight to your inbox weekly.