
Cloud Platform Engineer – Data Reliability, Backing Services
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in Egypt.
• Take ownership of the reliability, performance, scalability, and operational integrity of MySQL, PostgreSQL, MongoDB, Elasticsearch/OpenSearch, Kafka, StarRocks, ClickHouse, and similar platforms.
• Establish best practices for the utilization of transactional, document, search, messaging, and analytical platforms by product engineering teams.
• Design and uphold highly available, scalable, and resilient platform services, incorporating replication, backup, recovery, failover, and disaster recovery functionalities.
• Conduct capacity planning, performance tuning, workload assessments, upgrades, patching, and lifecycle management for platform services.
• Detect and mitigate risks including slow queries, hot partitions, consumer lag, replication lag, index growth, retention challenges, and storage saturation.
• Diagnose and resolve complex production issues associated with databases, messaging systems, search platforms, and distributed data platforms.
• Facilitate and support Kubernetes-based deployments of database, messaging, search, and analytical platforms using cloud-native methodologies and operational best practices.
• Leverage automation, Infrastructure as Code, GitOps, and CI/CD to ensure backing services are repeatable, reliable, and easier to manage.
• Collaborate with teams in Product Engineering, SRE, Security, Data Engineering, and Cloud Platform to enhance reliability, performance, availability, and security posture.
• 4-7 years of experience in Platform Engineering, SRE, Database Reliability Engineering, Data Platform Engineering, DevOps, or similar roles.
• Extensive hands-on experience with MySQL, PostgreSQL, Kafka, and at least one of MongoDB or Elasticsearch/OpenSearch in production settings.
• Familiarity with analytical or distributed data platforms such as StarRocks, ClickHouse, Apache Doris, Druid, Pinot, or comparable OLAP systems is highly preferred.
• Practical experience managing stateful workloads in Kubernetes-based environments.
• Solid understanding of high availability, replication, backup and recovery, disaster recovery, capacity planning, and performance tuning principles.
• Knowledge of distributed systems concepts including sharding, replication, partitioning, consistency, compaction, backpressure, consumer lag, and query optimization.
• Experience with at least one major cloud service provider (AWS, GCP, or OCI).
• Experience in automating provisioning, deployment, configuration, monitoring, and lifecycle management using tools such as Terraform, Helm, Ansible, GitOps, or similar automation frameworks.
• Strong scripting or programming abilities in Python, Bash, Go, or similar languages.
• Experience with observability platforms such as LTGM, Prometheus/Grafana, ELK/OpenSearch, Datadog, or equivalent solutions.
• Excellent troubleshooting, problem-solving, and debugging skills within distributed systems.
• Outstanding communication, collaboration, and documentation capabilities.
• Demonstrated curiosity, ownership mentality, adaptability, and the capacity to guide product engineering teams.
• Be part of an exhilarating phase for the Middle East, joining a rapidly growing company in an exciting industry.
• Enjoy significant responsibility and trust; we believe optimal results arise when individuals are given the freedom to execute their best judgment.
• Essential needs will be addressed: competitive salary, premium health insurance, and a supportive culture allowing you to concentrate on your strengths.
• Experience a vibrant and dynamic work environment alongside some of the brightest minds in AI.
• We value diversity, welcoming everyone for who they are and empowering them to become the best version of themselves.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.