
Distributed Systems Engineer III
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Egypt.
• Develop, manage, and enhance production cloud-native platforms that support distributed workloads.
• Engage directly with Kubernetes, focusing on upgrades, node pools, workload lifecycles, troubleshooting, and platform management.
• Implement and oversee workloads using ArgoCD, Helm, GitOps, Terraform, and automation techniques.
• Design and maintain systems emphasizing scalability, availability, resilience, performance, and operational efficiency.
• Diagnose intricate issues spanning Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services.
• Operate and resolve issues with Apache Kafka in production environments handling high-throughput and distributed workloads.
• Manage Kafka topics, partitions, replication, consumer groups, retention, throughput, latency, and recovery from failures.
• Connect Kafka with databases and applications through Kafka Connect, Debezium, or similar CDC/event-streaming solutions.
• Administer and troubleshoot MySQL and/or PostgreSQL, focusing on replication, high availability, backups, recovery, performance, and migrations.
• Support data-intensive workloads and analytical platforms like StarRocks, ClickHouse, Apache Doris, or similar technologies.
• Design and maintain resilient systems capable of withstanding node, service, zone, and infrastructure failures.
• Implement and verify backup, recovery, disaster recovery, failover, and business continuity capabilities.
• Apply principles of multi-zone, multi-region, and active-active architecture as appropriate.
• Construct multi-tenant platforms ensuring proper isolation, scalability, resource management, and reliability.
• Engage in disaster recovery exercises, failure simulations, migrations, and initiatives focused on resilience.
• Automate infrastructure and platform lifecycle operations using Terraform, Python, Bash, Go, or similar technologies.
• Establish reliable deployment and GitOps workflows to minimize manual operational efforts.
• Implement monitoring, logging, alerting, and observability for distributed workloads.
• Participate in production incident responses, root-cause analyses, and efforts for long-term reliability enhancements.
• Contribute infrastructure for AI, machine learning, and data-intensive workloads.
• Evolve cloud and Kubernetes platforms to accommodate AI workloads, data pipelines, model-serving infrastructure, and platform services.
• Address compute, GPU, networking, storage, data movement, observability, and workload isolation needs for AI platforms.
• Collaborate with engineering teams to create reusable platform capabilities for running AI and data workloads reliably at scale.
• Stay informed about infrastructure patterns across AI platforms, distributed data systems, and cloud-native technologies.
• 4–7 years of experience in Platform Engineering, Infrastructure Engineering, Distributed Systems, SRE, Backend Engineering, Data Infrastructure, or a related field.
• Strong production experience with Apache Kafka is essential.
• Hands-on production experience with Kubernetes is essential.
• Significant experience with at least one of MySQL or PostgreSQL is essential.
• Solid grasp of distributed systems fundamentals, including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery.
• Familiarity with ArgoCD/GitOps and infrastructure-as-code tools such as Terraform.
• Experience operating workloads on public cloud platforms like GCP, OCI, AWS, or Azure.
• Strong troubleshooting and incident resolution abilities in a production environment.
• Proficiency in automation or scripting using Python, Bash, Go, Java, or similar languages.
• Understanding of high availability, disaster recovery, and multi-tenant architecture fundamentals.
• Experience with observability and operational tools such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalent.
• Strong knowledge of infrastructure and networking principles in cloud-native environments.
• Good-to-have: experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms.
• Good-to-have: experience with distributed analytical databases like StarRocks, ClickHouse, Apache Doris, or similar.
• Good-to-have: experience with Flink, Spark, or other distributed data-processing technologies.
• Good-to-have: experience executing large-scale data, database, application, or infrastructure migrations.
• Good-to-have: experience with active-active, multi-zone, or multi-region systems.
• Good-to-have: experience managing stateful workloads on Kubernetes.
• Good-to-have: experience supporting AI/ML infrastructure or GPU-based workloads.
• Good-to-have: knowledge of cloud networking, service mesh, ingress, load balancing, or storage platforms.
• Good-to-have: contributions to Kubernetes, Kafka, distributed systems, or other open-source infrastructure projects.
• Competitive compensation.
• Top-tier health insurance.
• Supportive culture.
• Responsibility and trust.
• Freedom and autonomy in decision-making.
• Fun and dynamic workplace environment.
• Opportunity to collaborate with leading AI professionals.
• Inclusive and diverse workplace culture.
HighLevel
Coinbase
Tether.to
Netflix
Get handpicked remote jobs straight to your inbox weekly.