
Senior Software Engineer - Distributed Systems
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in New York.
β’ Design, develop, and scale distributed systems utilizing Kotlin, Java, and Spring Boot, while preparing for traffic spikes and concurrency challenges of a live real-money trading platform.
β’ Take ownership of the development and enhancement of high-throughput event pipelines using Kafka, focusing on partition strategies, consumer group architecture, and ensuring exactly-once processing guarantees.
β’ Construct and optimize low-latency data layers with Postgres, Redis, and Redis PubSub β maintaining data integrity, cache coherence, and achieving sub-millisecond read times during peak demand.
β’ Apply and uphold resilience patterns across services β including backpressure management, circuit breaking, idempotent retry mechanisms, and graceful degradation during failure scenarios.
β’ Engage in the development of real-time user-facing systems, tackling the fan-out challenge for live market updates across tens of thousands of concurrent sessions.
β’ Collaborate with product and engineering leadership to synchronize technical execution with business objectives β contributing to decisions regarding build vs. buy and shaping long-term platform strategy.
β’ Establish and maintain engineering standards for observability, schema evolution, testing methodologies, and deployment practices throughout the team β facilitated through code reviews, RFCs, and technical documentation.
β’ Mentor and actively cultivate junior and mid-level engineers, enhancing the technical capabilities of the team through collaborative work, design reviews, and direct feedback.
β’ 5+ years of experience in software engineering with a primary emphasis on distributed systems and high-concurrency production environments.
β’ Expert-level knowledge in Java or Kotlin and Spring Boot, complemented by a solid understanding of modern API design β including REST, gRPC, and Protobuf.
β’ Extensive hands-on experience with Kafka (or Redpanda/Pub Sub) β encompassing internal mechanics, partition strategies, consumer group rebalancing, and delivery guarantees.
β’ Proven capability to identify and resolve bottlenecks in asynchronous messaging systems and apply patterns such as idempotency, distributed caching, and exactly-once processing.
β’ Practical experience with Kubernetes, Helm, Terraform, and cloud-native infrastructure within AWS.
β’ Strong instincts for production reliability β experience being on-call, triaging distributed system failures under pressure, and delivering durable solutions.
β’ Demonstrated ability to influence technical direction across teams and guide engineers through intricate architectural decisions without direct authority.
β’ History of defining success metrics from the outset β including SLAs, latency budgets, and throughput targets β and ensuring systems are held accountable to these metrics in production.
β’ For information about our benefits, please visit https://benefitsatfanatics.com/.
Coinbase
DMS International
Netflix
Netflix
Get handpicked remote jobs straight to your inbox weekly.