
Software Engineer, Site Reliability
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Turkey.
• Take charge of our Kubernetes infrastructure, including cluster management, upgrades, networking, and ensuring multi-tenant isolation for customer workloads.
• Develop and sustain CI/CD pipelines as well as deployment infrastructure.
• Utilize AI to extensively automate the analysis and resolution of production issues, enhancing software development speed, reliability, and maintainability.
• Create dashboards, set up alerting mechanisms, and implement anomaly detection across our systems.
• Establish and enforce Service Level Objectives (SLOs) while developing incident response processes.
• Oversee and enhance our networking, load balancing, and service mesh configurations.
• Propel reliability enhancements throughout the stack via automation, runbooks, and chaos engineering practices.
• Minimum of 5 years of experience managing critical production systems and software development workflows.
• Extensive production experience in setting up and operating Kubernetes at scale, utilizing infrastructure-as-code tools such as Terraform and Ansible.
• In-depth understanding of Linux networking, container networking (including CNI plugins, VXLAN, BGP), and DNS.
• Proven experience in building CI/CD systems and implementing GitOps workflows (e.g., FluxCD, ArgoCD).
• Proficient in Python, as well as either Go or Bash for tooling and automation purposes.
• Strong background in logging, monitoring, and alerting tools (such as Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog).
• Exceptional communication skills with the ability to influence technical decisions across teams.
• A self-starter who executes efficiently, takes ownership, and continuously seeks improvement.
• Nice to have: Experience in managing GPU and AI/ML workloads.
• Nice to have: Familiarity with kernel-based monitoring and routing techniques (eBPF, XDP).
• Nice to have: Experience with security tools (e.g., Falco, Coroot, SIEM).
• Nice to have: Experience in bare metal Kubernetes networking (Calico, Cilium, MetalLB).
• Nice to have: Experience with distributed storage systems (such as Ceph, Longhorn, etc.).
• Engaging and challenging work opportunities.
• Abundant learning and growth prospects.
• Regular team events and offsite gatherings.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.