Senior Engineer, Cloud Infrastructure and Networking

atSkyloRemoteUS flagUnited StatesFull-timeCloud EngineerSenior$125k – $135k/year

Posted 14 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the 24x7 health of cloud infrastructure within Skylo’s hybrid production environment.

• Manage both GCP public cloud and on-premise private cloud infrastructures.

• Monitor and address infrastructure alarms utilizing OSS dashboards, Grafana/VictoriaMetrics, GCP Cloud Monitoring, and Loki.

• Implement runbooks for GKE node recovery, pod eviction/rescheduling, PVC repair, database failover, Prometheus WAL recovery, ArgoCD drift remediation, and certificate rotation.

• Oversee the observability pipeline, which includes Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, and alert routing.

• Maintain PostgreSQL replication, backups, restores, failover testing, query performance, and Redis operations.

• Ensure the health of the log aggregation pipeline using Loki or ELK.

• Collaborate with Network Implementation on infrastructure modifications and operational readiness.

• Act as the L3 escalation authority for Cloud Infrastructure incidents.

• Lead troubleshooting sessions and diagnose failures in Kubernetes, storage, network, database, and GitOps.

• Engage in global 24x7 on-call rotation responsibilities.

• Define and manage SLOs, monitor error budgets, minimize toil, and spearhead capacity planning.

• Conduct root-cause analyses for infrastructure issues and manage post-incident action items.

• Create and maintain Cloud Infrastructure runbooks and Standard Operating Procedures (SOPs).

• Validate operational readiness for infrastructure expansions, upgrades, and hardware deployments.

• Represent Cloud Infrastructure in architecture reviews and collaborate with NRE, security, platform engineering, and automation teams.

• Mentor Senior NREs in areas such as Kubernetes, storage, database reliability, observability, and escalation practices.

• Oversee changes involving ArgoCD, Helm, Terraform, and Ansible that affect production.


⛳️ Requirements

• Minimum of 5 years of experience in infrastructure engineering, Site Reliability Engineering, or cloud operations within a production 24x7 environment.

• Direct on-call ownership experience in Kubernetes-at-scale environments.

• Extensive knowledge of Kubernetes, including multi-cluster operations, node pool management, RBAC, network policies, PVCs, CSI drivers, CRD/operator patterns, and production cluster upgrades.

• Practical experience managing public cloud and on-premise/private cloud infrastructures.

• Ownership of production observability stacks utilizing Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, and Pub/Sub or similar tools.

• Familiarity with PostgreSQL streaming replication, backup/restore procedures, failover processes, and performance tuning.

• Experience managing Redis clusters and persistence.

• Background in production GitOps with ArgoCD or Flux CD.

• Proficient in Helm chart authorship and version management.

• Knowledge of Terraform or Ansible for infrastructure provisioning.

• Understanding of SRE fundamentals, including SLO/SLI/SLA definitions, error budget management, toil measurement, capacity planning, and on-call rotation design.

• Skills in container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.

• Ability to create runbooks that less-experienced engineers can execute independently under incident pressure.

• Strong written and verbal communication skills for RCA documents, structured engineering escalations, and MNO-facing infrastructure summaries.

• Preferred: Experience in telecom or NTN workloads.

• Preferred: Expertise in Ceph, Rook, or equivalent distributed storage technologies.

• Preferred: Familiarity with KubeVirt, Harvester, or OpenStack.

• Preferred: Knowledge of BGP, VXLAN, EVPN, software-defined networking, and hardware load balancers.

• Preferred: Experience in Go or Python development.

• Preferred: Background in FinOps.

• Preferred certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.

• Must be legally authorized to work in the U.S.


🏝️ Benefits

• Stock option-based equity program.

• Comprehensive medical, dental, and vision benefits.

• Retirement plan.

• Monthly wellness allowance.

• Monthly education reimbursement.

• Generous time-off policy.

• Paid holidays.

• Opportunity to work abroad temporarily.

• Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization.

• Flexible work approach.

• Inclusive and diverse workplace culture.

People also viewed

inexogy smart metering12 hours ago

Senior Cloud Platform Engineer

DE flagGermany OnlyFull-timeCloud Engineer
ApplyView job
SYNCREON16 hours ago

Lead Cloud Native Engineer, GCP

US flagNew Jersey OnlyFreelanceCloud Engineer
ApplyView job
TechPionier19 hours ago

IT Specialist, Cloud Computing

DE flagGermany OnlyFull-timeCloud Engineer€40k/year
ApplyView job
OpenText21 hours ago

Senior Cloud Services Project Manager

US flagUnited States OnlyFull-timeCloud Engineer
ApplyView job
group24 AG21 hours ago

Cloud Engineer – Azure

DE flagGermany OnlyFull-timeCloud Engineer
ApplyView job
MariaDB23 hours ago

Senior Software Engineer – Cloud Platform Engineering

US flagUnited States OnlyFull-timeCloud Engineer$141k – $185k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers