
Principal Engineer I, Prepurchase Platform
Posted Jun 19

Posted Jun 19
This is a fully remote position, open to applicants in Arizona, +3 more states.
• Proactively identify and address production issues without needing prompts. A single 500 error impacting a user should be treated as a critical issue, warranting immediate action regardless of overall traffic volume.
• Develop robust end-to-end systems with high availability that can manage significant traffic surges during peak sales events without performance degradation.
• Spearhead initiatives for platform modernization by transitioning legacy systems to cloud-native, microservice architectures.
• Create and deploy streaming and event-driven solutions utilizing Kafka/gRPC to facilitate real-time data flow between services.
• Immerse yourself in specific service areas when necessary or collaborate across teams to address cross-functional challenges.
• Advocate for resilience patterns such as circuit breakers, graceful degradation, bulkheads, and auto-scaling strategies.
• Work closely with the Enterprise Architecture team to influence architectural decisions and enhance engineering standards throughout the Prepurchase domain.
• Collaborate with product, security, and SRE teams to ensure technical choices align with business objectives.
• Lead efforts to enhance observability, ensuring services are properly instrumented for monitoring, alerting, and swift incident response.
• Identify and rectify single points of failure to bolster system reliability and lessen the on-call workload.
• Leverage AI and machine learning tools to boost developer productivity, automate operational tasks, and augment system functionalities.
• Assess and integrate new technologies that enhance performance, reliability, or engineering efficiency.
• Demonstrated ability to analyze the trade-offs of distributed systems, including scalability, consistency, availability, and latency, and make sound design choices under real-world constraints.
• Proven experience in writing production-grade code and constructing high-traffic systems at scale; this position prioritizes coding expertise.
• Proficient in Java (17+) and the JVM, with a strong understanding of JVM internals, garbage collection behavior, and performance optimization under load. Knowledge of Kotlin is a plus.
• Practical experience with a reactive, non-blocking framework (such as Vert.x, Spring WebFlux, or similar) and asynchronous, event-driven service architectures.
• Extensive experience with stream processing and event-driven architecture, such as Kafka Streams or Apache Flink, focusing on stateful streaming, topology design, and managing local state stores (e.g., RocksDB) at scale.
• A strong history of building high-throughput, low-latency systems where considerations for tail latency, backpressure, and memory pressure are integral to the design.
• Solid expertise in gRPC (including streaming RPC) and binary serialization formats like FlatBuffers, Avro, or Protobuf, with schema evolution managed via a registry.
• Understanding of Data-Oriented Design (DoD) principles, focusing on structuring code based on data layout, access patterns, and transformation for cache efficiency rather than object hierarchies.
• Experience with compact, search-optimized data structures (e.g., RoaringBitmap, succinct/bitset representations) for efficient large in-memory state representation.
• Familiarity with search engines (like Elasticsearch or Solr) for discovery workloads is an advantage.
• Strong grasp of microservice architecture, service mesh technologies (such as Istio/Envoy), API contract evolution, and backend-for-frontend patterns, including GraphQL APIs and WebSocket subscriptions for real-time client delivery.
• Practical experience with AWS (EKS) and cloud-native practices, including containerization (Docker, Kubernetes), packaging and deployment tools (Helm, Kustomize), and infrastructure-as-code (Terraform).
• Proficient in caching strategies (Redis, CDN caching) and their implementation in high-traffic systems.
• Comprehensive understanding of CI/CD pipelines (GitLab CI) and progressive deployment methods (such as blue-green and canary deployments).
• Experience in designing for high availability, including mult-region active-active/active-passive architectures, failover strategies, disaster recovery, and maintaining system stability under adverse conditions.
• Expertise in AI/ML tools and techniques, utilizing LLMs, AI-assisted development, and automation to enhance engineering processes and improve system intelligence.
• Familiarity with observability tools such as Grafana, Splunk, Prometheus, OpenTracing, or similar.
• Strong knowledge of security best practices, including OAuth/OIDC, input validation, and secrets management.
• Proven record of leading technical initiatives across multiple teams without direct authority.
• Comprehensive medical, vision, dental, and mental health benefits for you and your family, along with access to a health care concierge, and Flexible or Health Savings Accounts (FSA or HSA).
• Complimentary concert tickets, generous paid time off including holidays, sick days, and personal days.
• 401(k) plan with company match and a stock reimbursement program.
• New parent programs that include caregiver leave, as well as support for fertility, adoption, fostering, or surrogacy.
• Opportunities for career and skill development through the School of Live, tuition reimbursement, and student loan repayment initiatives.
• Volunteer time off and a crowdfunding match program.
Phase2
Job Mobz
RTX
Get handpicked remote jobs straight to your inbox weekly.