
Senior Site Reliability Engineer
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Egypt.
• Participate in a standard on-call rotation and serve as an active incident responder.
• Analyze, resolve, and document production incidents.
• Conduct application-level debugging and determine root causes within service code and business logic.
• Directly deliver fixes or pull requests into application repositories when necessary.
• Design, develop, and deploy LLM-based agents that are integrated with Kubernetes, cloud APIs, observability, and incident-management tools.
• Define interfaces for agent tools and create secure wrappers for APIs, scripts, and read/write operations.
• Set up autonomy guardrails and establish human-approval criteria for agents.
• Take ownership of agent evaluations and develop test/backtest suites using historical incident data.
• Optimize prompts, context, and tool schemas as the scope of agents expands.
• Collaborate with the SRE/platform team to discover suitable automation workflows.
• Measure agent impact through MTTD, MTTR, MTTX, false-positive/negative rates, and the reduction of engineer-hours of toil.
• Uphold security- and compliance-focused operations, ensuring audit trails, least-privilege access, and adherence to Saudi data-residency regulations.
• A minimum of 3 years of experience in building production software with LLMs, including agentic workflows, tool/function calling, multi-step planning, or RAG.
• Practical experience in delivering work using Claude Code, OpenAI Codex, or Kimi K2/K3.
• Proficiency in Python or similar programming languages for agent tooling, API wrappers, and orchestration.
• Practical SRE experience as a primary on-call responder, encompassing incident response and root cause analysis.
• Skills in application-level debugging along with the ability to interpret service code, trace failures to their underlying logic, and implement fixes.
• Hands-on experience with Kubernetes and cloud services such as AWS, GCP, OCI, or Azure.
• Familiarity with monitoring tools such as Prometheus, Grafana, Datadog, or ELK.
• Knowledge of guardrails for autonomous systems, permissioning, approval gates, rollback mechanisms, and auditability.
• Capability to build trust with technical stakeholders.
• Experience in Saudi Arabia/MENA, along with skills in Terraform/Ansible, Docker, VM/on-prem setups, LLM agent evaluation, ML engineering, LLMOps, or platform engineering, is a plus.
• Competitive compensation.
• Top-tier health insurance.
• Responsibility and trust.
• Freedom and autonomy in decision-making.
• A fun and dynamic workplace.
• An opportunity to collaborate with leading AI talent.
• An inclusive and empowering culture.
• The chance to work at the cutting edge of AI in the Middle East.
TEKsystems
Arctiq
GE Vernova
Ecosistemas
Get handpicked remote jobs straight to your inbox weekly.