
Senior SRE / DevOps Engineer, Kubernetes, A2P Messaging
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Poland.
• Utilize AI-enhanced engineering for tools related to infrastructure, automation, documentation, and runbooks.
• Build and manage Kubernetes infrastructure across two locations using infrastructure-as-code and automated delivery methods.
• Create deployment mechanisms that effectively drain long-standing connections during releases.
• Oversee platform networking, including reliable ingress and egress IPs, Layer 4 load balancing, TLS, and connectivity with external providers.
• Develop and maintain observability through metrics, logs, dashboards, and alerting systems.
• Design monitoring systems focused on production behavior, individual connections, and potential failure modes.
• Manage certificate lifecycles, secrets, and platform access controls.
• Operate highly available PostgreSQL solutions, encompassing replication, failover, backups, and verified restores.
• Design and implement backup and disaster-recovery procedures across two sites.
• Collaborate with engineers and QA teams on performance and production-scale load testing.
• Assist with migration and production cutovers.
• Engage in production incident response and participate in a 24×7 on-call rotation during normal operations.
• Build and manage a high-volume A2P messaging platform that processes approximately 1 million messages daily for around 175 customers and 120 suppliers, with over 300 long-lived messaging connections.
• Ensure the platform's progression through production readiness, migration, hypercare, and steady-state operation.
• Experience in telecom, carrier, messaging, or similar environments centered on persistent network connections.
• Familiarity with SMPP, SIP, SS7, or equivalent telecom protocols.
• Extensive experience in operating Kubernetes within self-managed, on-premise, or similarly infrastructure-intensive environments.
• Proficiency with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing, and deployment behavior.
• Solid understanding of Linux and networking fundamentals.
• Capability to diagnose routing, NAT, firewalls, MTU, TLS, and packet-level behaviors.
• Practical troubleshooting experience using tcpdump and production network diagnostics.
• Experience in designing monitoring and alerting systems.
• Strong background in infrastructure-as-code and CI/CD for containerized systems.
• Hands-on experience with PostgreSQL operations, including replication, failover, backup, and verified restoration.
• Ability to collaborate directly with engineers from external infrastructure and network providers.
• Proficient in professional English.
• Experience responding to real production incidents and being on-call.
• Familiarity with configuring or troubleshooting IPsec connectivity with external entities.
• Experience in designing or managing multi-site active-active or active-passive environments.
• Practical experience with disaster-recovery exercises.
• Background in performance engineering, including Linux kernel or network tuning for high connection counts.
• Experience in security hardening, such as CIS benchmarks, vulnerability management, or software supply-chain practices.
• Previous experience as the first SRE or Platform Engineer on a system.
• Fully remote work options available in Poland or the opportunity to work from the Łódź office.
• Autonomy and responsibility in your role.
• Commitment to diversity and psychological safety in the workplace.
• Agile and DevOps-oriented working environment.
• Opportunity to work on greenfield infrastructure with an influence over architecture and operational practices.
• Small senior team that promotes quick decision-making.
• Direct collaboration with engineers and the solution architect.
• AI-assisted tools that enhance engineering and operational tasks.
Truelogic Software
Cadwell
Raya
Arize AI
Get handpicked remote jobs straight to your inbox weekly.