
Staff Engineer – Distributed Systems
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Design, develop, and scale backend systems and frontend interfaces that enhance the Threads & Composer experience within HighLevel Conversations.
• Take ownership of backend and API design, data flows, performance, reliability, throughput, and accuracy at scale.
• Manage the Vue 3 UI layer comprehensively.
• Ensure the architectural integrity of a billion-scale distributed system, focusing on failure modes, capacity constraints, consistency guarantees, and interactions across over 50 deployments.
• Approve critical-path designs and influence architectural decisions across various teams.
• Proactively identify and resolve single points of failure, unbounded queues, absent idempotency, thundering herds, and data-loss intervals.
• Prototype high-risk architectural modifications, address significant incidents, and collaborate on complex cross-team bugs.
• Foster resilience through degradation strategies, backpressure, isolation boundaries, and capacity models that support over 10% month-over-month growth.
• Elevate engineering standards through design reviews, post-mortems, and the development of reusable patterns.
• Establish secure AI-assisted engineering practices for essential systems.
• Work with technologies such as Node.js/TypeScript, Go, GKE, GCP Pub/Sub, Cloud Tasks, Redis, MongoDB, Firestore, ClickHouse, and Elasticsearch.
• Enhance critical workflow paths through ownership, capacity models, tested failure modes, and incident reduction.
• Over 10 years of engineering experience with extensive, hands-on ownership of large-scale distributed systems.
• Sole responsibility for a production system through actual failures.
• Profound understanding of queuing and asynchronous architectures, delivery semantics, ordering, backpressure, idempotency, and both exactly-once and at-least-once delivery methods.
• Strong expertise in various SQL and NoSQL storage engines, consistency models, and indexing at scale.
• Expert-level knowledge in Redis or similar in-memory systems, including failure modes under memory constraints and network partitions.
• Production experience with Kubernetes at scale, including resource limits, autoscaling behaviors, and node-pool failures.
• Exceptional design communication skills through documentation, diagrams, and root-cause analyses.
• Proficiency in Node.js and/or Go sufficient to prototype proposals and implement critical-path fixes.
• Comparable large-scale experience with hundreds of services or thousands of instances handling billions of daily events is strongly preferred.
• Experience with AI agents, GCP-native services, transitioning 0→1 systems to production, and fortifying mature systems is considered a bonus.
• Global, remote-first organization.
• Equal Opportunity Employer.
• Voluntary demographic information process; this information is kept separate from applications and not utilized in hiring decisions.
• Privacy Policy available for review.
Coinbase
Tether.to
Netflix
Get handpicked remote jobs straight to your inbox weekly.