
Site Reliability Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Design and develop cloud platforms that support OXIO backend services.
• Automate technical operations, which encompass deployments, scaling, and recovery processes.
• Monitor and uphold mission-critical production infrastructure to ensure maximum uptime.
• Engage in an on-call rotation and foster continuous improvement through blameless postmortems.
• Equip Engineering, Telecom, and Data Engineering teams with tools that facilitate their service operations.
• Contribute to the development of OXIO’s Carrier-as-a-Service telecom platform and modern connectivity infrastructure.
• Comprehensive understanding of Linux/Unix systems.
• Familiarity with Linux/Unix system internals, including process management, filesystems, memory management, and networking.
• Proficient in at least one programming language: Python, Go, or Ruby.
• Strong scripting capabilities in Bash or Perl.
• Experience with infrastructure provisioning tools like Terraform, CloudFormation, or Ansible.
• Knowledge of Docker and Kubernetes.
• Experience with monitoring tools such as Prometheus, Grafana, or Datadog.
• Understanding of alerts, log analysis, dashboards, and observability practices.
• Familiarity with incident management methodologies, including runbooks and postmortems.
• Experience in participating in an on-call rotation and effectively managing incidents.
• Proficient in setting up and maintaining CI/CD pipelines like Jenkins, GitLab CI, or CircleCI.
• Hands-on experience with cloud platforms such as AWS, Google Cloud, or Azure.
• Knowledge of VMware, KVM, and cloud-native architecture.
• Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewall configurations.
• Nice-to-have: experience with deployment strategies, high availability and failover, IAM and zero trust principles, distributed systems, custom monitoring, SQL and NoSQL databases, distributed tracing, log aggregation, performance profiling, load testing, and SaltStack configuration management.
• Competitive salary and performance-based incentives.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Professional development opportunities and support for further education.
• Collaborative and inclusive work environment.
General Dynamics Information Technology
VALCE Talent Solutions
VALCE Talent Solutions
ImmunityBio, Inc.
Get handpicked remote jobs straight to your inbox weekly.