
Senior Engineer, Cloud Infrastructure and Networking
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the 24x7 health of cloud infrastructure within Skylo’s hybrid production environment.
• Manage both GCP public cloud and on-premise private cloud infrastructures.
• Monitor and address infrastructure alarms utilizing OSS dashboards, Grafana/VictoriaMetrics, GCP Cloud Monitoring, and Loki.
• Implement runbooks for GKE node recovery, pod eviction/rescheduling, PVC repair, database failover, Prometheus WAL recovery, ArgoCD drift remediation, and certificate rotation.
• Oversee the observability pipeline, which includes Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, and alert routing.
• Maintain PostgreSQL replication, backups, restores, failover testing, query performance, and Redis operations.
• Ensure the health of the log aggregation pipeline using Loki or ELK.
• Collaborate with Network Implementation on infrastructure modifications and operational readiness.
• Act as the L3 escalation authority for Cloud Infrastructure incidents.
• Lead troubleshooting sessions and diagnose failures in Kubernetes, storage, network, database, and GitOps.
• Engage in global 24x7 on-call rotation responsibilities.
• Define and manage SLOs, monitor error budgets, minimize toil, and spearhead capacity planning.
• Conduct root-cause analyses for infrastructure issues and manage post-incident action items.
• Create and maintain Cloud Infrastructure runbooks and Standard Operating Procedures (SOPs).
• Validate operational readiness for infrastructure expansions, upgrades, and hardware deployments.
• Represent Cloud Infrastructure in architecture reviews and collaborate with NRE, security, platform engineering, and automation teams.
• Mentor Senior NREs in areas such as Kubernetes, storage, database reliability, observability, and escalation practices.
• Oversee changes involving ArgoCD, Helm, Terraform, and Ansible that affect production.
• Minimum of 5 years of experience in infrastructure engineering, Site Reliability Engineering, or cloud operations within a production 24x7 environment.
• Direct on-call ownership experience in Kubernetes-at-scale environments.
• Extensive knowledge of Kubernetes, including multi-cluster operations, node pool management, RBAC, network policies, PVCs, CSI drivers, CRD/operator patterns, and production cluster upgrades.
• Practical experience managing public cloud and on-premise/private cloud infrastructures.
• Ownership of production observability stacks utilizing Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, and Pub/Sub or similar tools.
• Familiarity with PostgreSQL streaming replication, backup/restore procedures, failover processes, and performance tuning.
• Experience managing Redis clusters and persistence.
• Background in production GitOps with ArgoCD or Flux CD.
• Proficient in Helm chart authorship and version management.
• Knowledge of Terraform or Ansible for infrastructure provisioning.
• Understanding of SRE fundamentals, including SLO/SLI/SLA definitions, error budget management, toil measurement, capacity planning, and on-call rotation design.
• Skills in container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.
• Ability to create runbooks that less-experienced engineers can execute independently under incident pressure.
• Strong written and verbal communication skills for RCA documents, structured engineering escalations, and MNO-facing infrastructure summaries.
• Preferred: Experience in telecom or NTN workloads.
• Preferred: Expertise in Ceph, Rook, or equivalent distributed storage technologies.
• Preferred: Familiarity with KubeVirt, Harvester, or OpenStack.
• Preferred: Knowledge of BGP, VXLAN, EVPN, software-defined networking, and hardware load balancers.
• Preferred: Experience in Go or Python development.
• Preferred: Background in FinOps.
• Preferred certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.
• Must be legally authorized to work in the U.S.
• Stock option-based equity program.
• Comprehensive medical, dental, and vision benefits.
• Retirement plan.
• Monthly wellness allowance.
• Monthly education reimbursement.
• Generous time-off policy.
• Paid holidays.
• Opportunity to work abroad temporarily.
• Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization.
• Flexible work approach.
• Inclusive and diverse workplace culture.
inexogy smart metering
SYNCREON
TechPionier
OpenText
Get handpicked remote jobs straight to your inbox weekly.