
Senior DevOps Engineer – PostgreSQL, Kafka, Kubernetes
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Deploy, upgrade, scale, and manage CloudNativePG PostgreSQL and Strimzi Kafka clusters on both cloud and bare-metal Kubernetes environments.
• Handle infrastructure management declaratively using Argo CD or Flux, Helm, and Terraform.
• Develop CI/CD pipelines for changes to databases and Kafka, schema migrations, and operator upgrades.
• Design and maintain regional high availability and cross-region disaster recovery strategies.
• Manage PostgreSQL replica clusters, backups, point-in-time recovery, and Kafka MirrorMaker 2 operations.
• Conduct regular failover drills with specified RPO/RTO targets.
• Enable product teams to provision databases, users, topics, and ACLs through self-service code solutions.
• Create monitoring and alerting systems using Prometheus, Grafana, and OpenTelemetry.
• Monitor replication lag, backup statuses, consumer lag, and overall capacity.
• Participate in on-call rotations and lead incident review sessions.
• Implement TLS/mTLS, manage secrets, enforce network policies, apply Kafka ACLs, and ensure per-tenant isolation.
• Integrate database and Kafka access with Keycloak/OIDC identity platforms.
• Operate Debezium and Kafka Connect outbox/change-data-capture pipelines.
• Develop automation and small services in Go or Python.
• Collaborate with application teams to establish schema standards and address performance issues.
• Over 8 years of experience in DevOps, SRE, platform, or infrastructure engineering roles.
• At least 3 years of experience managing stateful systems (such as databases or message streaming) in a production environment.
• Strong expertise in Kubernetes operations, including StatefulSets, persistent volumes and storage classes, pod disruption budgets, scheduling and affinity, network policies, and cluster upgrades.
• Practical experience running PostgreSQL and/or Kafka via Kubernetes operators in production.
• Proficiency in authoring and maintaining Helm charts.
• Daily usage of Argo CD or Flux, Terraform, and CI/CD pipelines.
• Hands-on PostgreSQL administration skills, covering replication and failover, backup and point-in-time recovery, connection pooling using PgBouncer, version upgrades, and performance troubleshooting.
• Experience in production Kafka operations, including brokers, topics and partitions, replication, consumer groups, monitoring, and capacity planning.
• Proven ability to design and test high availability and disaster recovery solutions for stateful systems, with defined RPO/RTO parameters.
• Strong scripting and programming capabilities in Go or Python, along with Bash.
• A Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
• Additional experience in cross-region replication, Debezium, Kafka Connect, transactional outbox/CDC patterns, Keycloak, OIDC, LDAP/Active Directory, observability stacks, Vault/OpenBao, cert-manager, mTLS, bare-metal infrastructure, Ceph or object storage, GPU/AI infrastructure, open-source contributions, and SOC 2 or ISO 27001 environments is a plus.
• Opportunities for professional development and training.
• Attendance at conferences and participation in working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package along with a comprehensive benefits plan.
• Flexible remote work arrangement.
Slate Auto
Funding Xchange
Leidos
LeoLabs
Get handpicked remote jobs straight to your inbox weekly.