
Senior DevOps Engineer – GCP
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Cyprus, +5 more countries.
• Design and manage GoMining's cloud platform.
• Build and maintain infrastructure on GCP, which includes GKE, virtual machines, networking, load balancers, storage, and various managed services.
• Oversee Infrastructure as Code with Terraform, focusing on modules, state management, dependencies, reproducible changes, and configuration drift.
• Manage Kubernetes operations by upgrading clusters and platform components, setting up autoscaling, workload placement, network policies, and resource allocation.
• Create CI/CD and GitOps workflows using GitLab CI and Flux / Argo CD to automate builds, validations, deployments, controlled releases, and rollbacks.
• Establish observability through metrics, logs, distributed tracing, and actionable alerting; collaborate with engineering teams to define and enhance SLIs/SLOs.
• Diagnose production incidents and performance challenges, perform root cause analyses, and mitigate recurring failure patterns.
• Enforce least-privilege access, ensure environment and service isolation, manage secret security and rotation, and implement infrastructure audit controls.
• Oversee backup processes and confirm the recovery of essential data and infrastructure in accordance with agreed RPO/RTO standards.
• Assist engineering teams in preparing services for production, including health checks, resource configurations, telemetry, and deployment protocols.
• Optimize cloud infrastructure usage and costs while upholding reliability standards.
• Maintain comprehensive infrastructure documentation, operational runbooks, and recovery procedures.
• Extensive hands-on experience in independently managing production infrastructure and Kubernetes, including upgrades, troubleshooting, and disaster recovery.
• Strong expertise in GCP, encompassing GKE, Compute Engine, VPC, IAM, and Cloud Storage; solid understanding of Cloud SQL, Cloud Logging, and Cloud Monitoring.
• Practical experience with Terraform, including modules, remote state management, locking, importing existing resources, and safely applying infrastructure changes via CI.
• Profound knowledge of Kubernetes and containerization: Deployments, StatefulSets, Services, Ingress, storage, requests/limits, probes, autoscaling, RBAC, and NetworkPolicy.
• Familiarity with Helm and Kustomize.
• Strong Linux administration capabilities, ideally with Debian/Ubuntu: processes, systemd, filesystems, disks, and resource troubleshooting.
• Solid understanding of networking fundamentals: TCP/IP, DNS, HTTP(S), TLS, routing, NAT, load balancing, and firewalls.
• Proficient in troubleshooting connectivity issues across applications, Kubernetes clusters, networks, and cloud services.
• Experience in constructing CI/CD pipelines, preferably utilizing GitLab CI, with a robust understanding of Git and GitOps principles.
• Knowledge of runner permissions, secrets protection, and production deployment controls.
• Automation skills in Bash and Python, focusing on maintainable infrastructure scripts and API integrations.
• Experience with monitoring, logging, and alerting using tools like Prometheus, Grafana, or Cloud Monitoring.
• Practical experience with PostgreSQL operations and familiarity with Redis and RabbitMQ, covering availability, connections, replication or clustering, backup, and recovery.
• In-depth understanding of cloud security principles: least privilege, service accounts, Workload Identity, secrets management, network segmentation, and access auditing.
• Capability to independently implement infrastructure changes leading to validated production outcomes, articulate technical trade-offs, and collaborate effectively with Engineering and Security teams.
• Nice to Have: Familiarity with Flux CD, Argo CD, Ansible, and HashiCorp Vault.
• Nice to Have: Experience with OpenTelemetry and distributed tracing.
• Nice to Have: Experience in migrating infrastructure between clouds or Kubernetes clusters, including database migrations with minimal downtime.
• Nice to Have: Skills in designing disaster recovery strategies and conducting recovery exercises.
• Nice to Have: Background in cloud cost optimization and capacity planning.
• Nice to Have: Experience with DigitalOcean, Hetzner, or AWS.
• Nice to Have: Ability to read and troubleshoot JavaScript / TypeScript applications.
• Nice to Have: Experience with operating financial or payment services that require strict access control, auditing, reliability, and data integrity standards.
• Professional development opportunities: support for courses, conferences, and English learning (up to 100% coverage).
• Work-life balance: remote or hybrid work options with flexible hours across international teams.
• Paid time off: up to 20 vacation days, plus 8 company holidays and 5 personal days each year.
• Recognition programs: structured performance reviews and team awards.
• Team culture: retreats in international locations (such as company apartments in Cyprus).
Akamai Technologies
BeyondTrust
Cencora
PrizePicks
Get handpicked remote jobs straight to your inbox weekly.