Senior SRE / DevOps Engineer, Kubernetes, A2P Messaging

Posted 4 days ago

This is a fully remote position, open to applicants in Poland.

📋 Description

• Utilize AI-enhanced engineering for tools related to infrastructure, automation, documentation, and runbooks.

• Build and manage Kubernetes infrastructure across two locations using infrastructure-as-code and automated delivery methods.

• Create deployment mechanisms that effectively drain long-standing connections during releases.

• Oversee platform networking, including reliable ingress and egress IPs, Layer 4 load balancing, TLS, and connectivity with external providers.

• Develop and maintain observability through metrics, logs, dashboards, and alerting systems.

• Design monitoring systems focused on production behavior, individual connections, and potential failure modes.

• Manage certificate lifecycles, secrets, and platform access controls.

• Operate highly available PostgreSQL solutions, encompassing replication, failover, backups, and verified restores.

• Design and implement backup and disaster-recovery procedures across two sites.

• Collaborate with engineers and QA teams on performance and production-scale load testing.

• Assist with migration and production cutovers.

• Engage in production incident response and participate in a 24×7 on-call rotation during normal operations.

• Build and manage a high-volume A2P messaging platform that processes approximately 1 million messages daily for around 175 customers and 120 suppliers, with over 300 long-lived messaging connections.

• Ensure the platform's progression through production readiness, migration, hypercare, and steady-state operation.


⛳️ Requirements

• Experience in telecom, carrier, messaging, or similar environments centered on persistent network connections.

• Familiarity with SMPP, SIP, SS7, or equivalent telecom protocols.

• Extensive experience in operating Kubernetes within self-managed, on-premise, or similarly infrastructure-intensive environments.

• Proficiency with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing, and deployment behavior.

• Solid understanding of Linux and networking fundamentals.

• Capability to diagnose routing, NAT, firewalls, MTU, TLS, and packet-level behaviors.

• Practical troubleshooting experience using tcpdump and production network diagnostics.

• Experience in designing monitoring and alerting systems.

• Strong background in infrastructure-as-code and CI/CD for containerized systems.

• Hands-on experience with PostgreSQL operations, including replication, failover, backup, and verified restoration.

• Ability to collaborate directly with engineers from external infrastructure and network providers.

• Proficient in professional English.

• Experience responding to real production incidents and being on-call.

• Familiarity with configuring or troubleshooting IPsec connectivity with external entities.

• Experience in designing or managing multi-site active-active or active-passive environments.

• Practical experience with disaster-recovery exercises.

• Background in performance engineering, including Linux kernel or network tuning for high connection counts.

• Experience in security hardening, such as CIS benchmarks, vulnerability management, or software supply-chain practices.

• Previous experience as the first SRE or Platform Engineer on a system.


🏝️ Benefits

• Fully remote work options available in Poland or the opportunity to work from the Łódź office.

• Autonomy and responsibility in your role.

• Commitment to diversity and psychological safety in the workplace.

• Agile and DevOps-oriented working environment.

• Opportunity to work on greenfield infrastructure with an influence over architecture and operational practices.

• Small senior team that promotes quick decision-making.

• Direct collaboration with engineers and the solution architect.

• AI-assisted tools that enhance engineering and operational tasks.

People also viewed

Truelogic Software15 hours ago

Senior DevOps Engineer – Wealth Management Fintech

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cadwell17 hours ago

Cloud Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $130k/year
ApplyView job
Raya18 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Arize AI19 hours ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $185k/year
ApplyView job
Capgemini22 hours ago

DevOps Engineer

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cast & Crew22 hours ago

Staff DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$190k – $235k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers