
Principal Backend Engineer, Performance, Database
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Pakistan.
• Take charge of the performance and scalability aspects of backend systems and their associated data layers.
• Craft request flows, differentiating between synchronous and asynchronous processes, state placement, dependency management, and establishing scalable service boundaries.
• Troubleshoot latency and performance challenges by utilizing traces, execution plans, metrics, and supporting evidence.
• Set up instrumentation, review protocols, and regression checkpoints to avert issues visible to customers.
• Devise strategies for read and write scaling, separation of read and write operations, background processing, real-time updates, handling traffic surges, ensuring high availability and reliability, managing data growth, implementing safe changes, and optimizing API performance.
• Evaluate architectural proposals from feature teams.
• Draft design documents and decision records to guide other engineers in their development efforts.
• Choose and refine relational, document, key-value, and in-memory data stores based on workload demands and access patterns.
• Analyze execution plans, index costs, lock contention, connection pooling, replica lag, partitioning, document aggregation, shard keys, hot partitions, cache behavior, and dual-write challenges.
• Investigate latency across the entire request path, including ORM-generated traffic, synchronous and ETL latency, as well as third-party integrations.
• Manage data-layer capacity planning and cost per request while monitoring latency.
• Introduce design and query evaluations for critical pathways and changes to the data layer.
• Establish standards for migrations, validate index necessity, and identify work that should not run synchronously.
• Integrate data access instrumentation within application traces.
• Contribute to CI regression checkpoints and conduct production-like load testing.
• Analyze performance for major customer accounts and tenant-specific behavior.
• Participate in the escalation on-call rotation once it's established.
• Over 5 years of experience in building and managing production backend systems.
• At least 3 years where performance and scalability were key responsibilities.
• Proven experience in managing performance for a multi-tenant SaaS product under real-world conditions.
• Demonstrated ability to design systems, showcasing rejected alternatives, accepted trade-offs, and lessons learned.
• Specific examples of diagnosing performance issues, pinpointing their causes, implementing solutions, and measuring the outcomes.
• Strong analytical skills regarding caching, replicas, indexes, invalidation, staleness, sharding, partitioning, batching, queue buffering, replica routing, CQRS, queues, workers, orchestration, idempotency, dead letters, WebSockets, SSE, long polling, load balancing, autoscaling, headroom, load shedding, replication, failover, degraded operations, timeouts, backoff, circuit breakers, dedicated search indexes, archival, retention, canary deployments, feature flags, online migrations, logs, metrics, traces, and alerts.
• Capability to create design documents that are actionable for teams and understandable for executives.
• Extensive experience with at least two types of databases: relational, document, key-value, and in-memory, with functional knowledge of the others.
• Proficiency in relational databases, including execution plans, composite and partial indexes, index write costs, locking mechanisms, isolation levels, connection pooling, and replica lag.
• Expertise in document databases, focusing on embedding versus referencing, shard-key selection, aggregation performance, and working-set sizing.
• Key-value database proficiency, including access-pattern modeling, partition key selection, hot-partition avoidance, secondary indexing, and capacity cost management.
• Cache management skills, including understanding caching patterns, eviction strategies, TTL, and behavior under cold-cache or cache-loss scenarios.
• Knowledge of consistency boundaries and the ability to identify when a query issue stems from application design flaws.
• Experience in production ownership across multiple backend programming languages.
• Understanding of ORM and ODM query generation processes.
• Familiarity with concurrency, connection, and resource management during high-load situations.
• Background in queue design, particularly in managing consumers that may lag behind.
• Experience utilizing APM and tracing tools for performance diagnosis.
• Hands-on experience with major public cloud services, including understanding managed data-service cost behaviors.
• Ability to instrument code for future diagnostic purposes.
• Full-time, permanent position.
• Flexible remote work options.
• High degree of autonomy and considerable technical influence.
• Opportunity to shape backend architecture and performance criteria.
• Potential for future leadership roles as the function expands and if the candidate aspires to lead.
Oscilar
Veradigm®
Get handpicked remote jobs straight to your inbox weekly.