
Principal Engineer I, Prepurchase Platform
Posted Jul 15

Posted Jul 15
This is a fully remote position, open to applicants in United Kingdom.
• Proactively identify and address production issues without external prompting.
• Develop end-to-end, high-availability systems capable of managing significant traffic surges during peak demand periods without performance loss.
• Spearhead platform modernization efforts by transitioning legacy services to cloud-native, microservices architectures.
• Design and implement streaming and event-driven solutions utilizing Kafka/gRPC to facilitate real-time data flow across various services.
• Engage deeply in a specific service area as needed, or collaborate across teams to tackle cross-functional challenges.
• Advocate for resilience patterns such as circuit breakers, graceful degradation, bulkheads, and auto-scaling strategies.
• Work closely with the Enterprise Architecture team on architectural decisions and enhance engineering standards within the Prepurchase domain.
• Collaborate with product, security, and SRE teams to ensure alignment of technical decisions with business objectives.
• Lead initiatives to improve observability by ensuring services are equipped for monitoring, alerting, and swift incident response.
• Identify and eliminate single points of failure to enhance system reliability and reduce on-call responsibilities.
• Utilize AI and machine learning tools to boost developer productivity, automate operational tasks, and enhance system capabilities.
• Assess and introduce new technologies that enhance performance, reliability, or engineering velocity.
• Demonstrated ability to analyze tradeoffs in distributed systems—scalability, consistency, availability, and latency—and make well-informed design decisions under real-world constraints.
• Proven experience in writing production-level code and developing high-traffic systems at scale; this role prioritizes coding expertise.
• Proficiency in Java (version 17+) and a strong understanding of JVM internals, garbage collection behavior, and performance tuning under load.
• Hands-on experience with reactive, non-blocking frameworks (such as Vert.x, Spring WebFlux, or similar) and designing asynchronous, event-driven services.
• Extensive experience with stream processing and event-driven architectures, including Kafka Streams and Apache Flink, with a focus on stateful streaming and topology design.
• Proven track record of building high-throughput, low-latency systems where tail latency, backpressure, and memory pressure are critical design considerations.
• Strong knowledge of gRPC (including streaming RPC) and binary serialization formats like FlatBuffers, Avro, or Protobuf, with schema evolution managed through a registry.
• Familiarity with Data-Oriented Design (DoD), focusing on structuring code based on data layout and access patterns for improved cache efficiency and throughput.
• Experience with compact, search-optimized data structures (e.g., RoaringBitmap, succinct/bitset representations) for efficient representation of large in-memory states.
• Experience with search engines (like Elasticsearch or Solr) for discovery workloads is advantageous.
• In-depth understanding of microservice design, service mesh (such as Istio/Envoy), API contract evolution, and backend-for-frontend patterns, including GraphQL APIs and WebSocket subscriptions for real-time client delivery.
• Practical experience with AWS (EKS) and cloud-native operations, including containerization (Docker, Kubernetes), deployment tooling (Helm, Kustomize), and infrastructure-as-code (Terraform).
• Proficiency in caching strategies (Redis, CDN layer caching) and their implementation in high-traffic systems.
• Solid understanding of CI/CD pipelines (GitLab CI) and progressive deployment strategies (blue-green, canary).
• Experience designing for high availability, including multi-region active-active/active-passive architectures, failover strategies, disaster recovery, and systems that maintain stability under adversarial load and abuse.
• Familiarity with observability tools such as Grafana, Splunk, Prometheus, OpenTracing, or similar.
• Strong understanding of security best practices, including OAuth/OIDC, input validation, and secrets management.
• Proven ability to lead technical initiatives across multiple teams without direct authority.
• Health insurance
• Retirement plans
• Professional development opportunities
• Paid time off
Phase2
Job Mobz
RTX
Get handpicked remote jobs straight to your inbox weekly.