Site Reliability Engineer

Posted 11 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Design and develop cloud platforms that support OXIO backend services.

• Automate technical operations, which encompass deployments, scaling, and recovery processes.

• Monitor and uphold mission-critical production infrastructure to ensure maximum uptime.

• Engage in an on-call rotation and foster continuous improvement through blameless postmortems.

• Equip Engineering, Telecom, and Data Engineering teams with tools that facilitate their service operations.

• Contribute to the development of OXIO’s Carrier-as-a-Service telecom platform and modern connectivity infrastructure.


⛳️ Requirements

• Comprehensive understanding of Linux/Unix systems.

• Familiarity with Linux/Unix system internals, including process management, filesystems, memory management, and networking.

• Proficient in at least one programming language: Python, Go, or Ruby.

• Strong scripting capabilities in Bash or Perl.

• Experience with infrastructure provisioning tools like Terraform, CloudFormation, or Ansible.

• Knowledge of Docker and Kubernetes.

• Experience with monitoring tools such as Prometheus, Grafana, or Datadog.

• Understanding of alerts, log analysis, dashboards, and observability practices.

• Familiarity with incident management methodologies, including runbooks and postmortems.

• Experience in participating in an on-call rotation and effectively managing incidents.

• Proficient in setting up and maintaining CI/CD pipelines like Jenkins, GitLab CI, or CircleCI.

• Hands-on experience with cloud platforms such as AWS, Google Cloud, or Azure.

• Knowledge of VMware, KVM, and cloud-native architecture.

• Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewall configurations.

• Nice-to-have: experience with deployment strategies, high availability and failover, IAM and zero trust principles, distributed systems, custom monitoring, SQL and NoSQL databases, distributed tracing, log aggregation, performance profiling, load testing, and SaltStack configuration management.


🏝️ Benefits

• Competitive salary and performance-based incentives.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Professional development opportunities and support for further education.

• Collaborative and inclusive work environment.

People also viewed

General Dynamics Information Technology7 hours ago

Principal DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$144.5k – $195.5k/year
ApplyView job
VALCE Talent Solutions7 hours ago

DevOps/SRE Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
VALCE Talent Solutions10 hours ago

Senior Release Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ImmunityBio, Inc.10 hours ago

DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130.5k – $150k/year
ApplyView job
Rimutee13 hours ago

DevOps Engineer – LATAM

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Rimutee14 hours ago

DevOps Engineer – LATAM, Español

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers