Remotery

Software Engineer, Site Reliability

Posted Jul 28

This is a fully remote position, open to applicants in Turkey.

📋 Description

• Take charge of our Kubernetes infrastructure, including cluster management, upgrades, networking, and ensuring multi-tenant isolation for customer workloads.

• Develop and sustain CI/CD pipelines as well as deployment infrastructure.

• Utilize AI to extensively automate the analysis and resolution of production issues, enhancing software development speed, reliability, and maintainability.

• Create dashboards, set up alerting mechanisms, and implement anomaly detection across our systems.

• Establish and enforce Service Level Objectives (SLOs) while developing incident response processes.

• Oversee and enhance our networking, load balancing, and service mesh configurations.

• Propel reliability enhancements throughout the stack via automation, runbooks, and chaos engineering practices.


⛳️ Requirements

• Minimum of 5 years of experience managing critical production systems and software development workflows.

• Extensive production experience in setting up and operating Kubernetes at scale, utilizing infrastructure-as-code tools such as Terraform and Ansible.

• In-depth understanding of Linux networking, container networking (including CNI plugins, VXLAN, BGP), and DNS.

• Proven experience in building CI/CD systems and implementing GitOps workflows (e.g., FluxCD, ArgoCD).

• Proficient in Python, as well as either Go or Bash for tooling and automation purposes.

• Strong background in logging, monitoring, and alerting tools (such as Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog).

• Exceptional communication skills with the ability to influence technical decisions across teams.

• A self-starter who executes efficiently, takes ownership, and continuously seeks improvement.

• Nice to have: Experience in managing GPU and AI/ML workloads.

• Nice to have: Familiarity with kernel-based monitoring and routing techniques (eBPF, XDP).

• Nice to have: Experience with security tools (e.g., Falco, Coroot, SIEM).

• Nice to have: Experience in bare metal Kubernetes networking (Calico, Cilium, MetalLB).

• Nice to have: Experience with distributed storage systems (such as Ceph, Longhorn, etc.).


🏝️ Benefits

• Engaging and challenging work opportunities.

• Abundant learning and growth prospects.

• Regular team events and offsite gatherings.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers