Senior Site Reliability Engineer

Posted Aug 18

This is a fully remote position, open to applicants in Egypt.

📋 Description

• Participate in a standard on-call rotation and serve as an active incident responder.

• Analyze, resolve, and document production incidents.

• Conduct application-level debugging and determine root causes within service code and business logic.

• Directly deliver fixes or pull requests into application repositories when necessary.

• Design, develop, and deploy LLM-based agents that are integrated with Kubernetes, cloud APIs, observability, and incident-management tools.

• Define interfaces for agent tools and create secure wrappers for APIs, scripts, and read/write operations.

• Set up autonomy guardrails and establish human-approval criteria for agents.

• Take ownership of agent evaluations and develop test/backtest suites using historical incident data.

• Optimize prompts, context, and tool schemas as the scope of agents expands.

• Collaborate with the SRE/platform team to discover suitable automation workflows.

• Measure agent impact through MTTD, MTTR, MTTX, false-positive/negative rates, and the reduction of engineer-hours of toil.

• Uphold security- and compliance-focused operations, ensuring audit trails, least-privilege access, and adherence to Saudi data-residency regulations.


⛳️ Requirements

• A minimum of 3 years of experience in building production software with LLMs, including agentic workflows, tool/function calling, multi-step planning, or RAG.

• Practical experience in delivering work using Claude Code, OpenAI Codex, or Kimi K2/K3.

• Proficiency in Python or similar programming languages for agent tooling, API wrappers, and orchestration.

• Practical SRE experience as a primary on-call responder, encompassing incident response and root cause analysis.

• Skills in application-level debugging along with the ability to interpret service code, trace failures to their underlying logic, and implement fixes.

• Hands-on experience with Kubernetes and cloud services such as AWS, GCP, OCI, or Azure.

• Familiarity with monitoring tools such as Prometheus, Grafana, Datadog, or ELK.

• Knowledge of guardrails for autonomous systems, permissioning, approval gates, rollback mechanisms, and auditability.

• Capability to build trust with technical stakeholders.

• Experience in Saudi Arabia/MENA, along with skills in Terraform/Ansible, Docker, VM/on-prem setups, LLM agent evaluation, ML engineering, LLMOps, or platform engineering, is a plus.


🏝️ Benefits

• Competitive compensation.

• Top-tier health insurance.

• Responsibility and trust.

• Freedom and autonomy in decision-making.

• A fun and dynamic workplace.

• An opportunity to collaborate with leading AI talent.

• An inclusive and empowering culture.

• The chance to work at the cutting edge of AI in the Middle East.

People also viewed

TEKsystems1 day ago

Cloud Deployment Engineer, Secret Clearance Required – 30% Travel

US flagDistrict of Columbia, +1 more stateFreelanceDevOps & Site Reliability Engineer (SRE)$80 – $110/hour
ApplyView job
Arctiq1 day ago

Site Reliability Engineer – Vulnerability Remediation Consultant

US flagUnited States OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
GE Vernova1 day ago

Senior Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$152.4k – $254k/year
ApplyView job
Ecosistemas1 day ago

Senior DevOps Engineer, Bilingüe Inglés

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Intetics1 day ago

Senior DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
OZmap1 day ago

Mid-Level SRE

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers