
Senior Elasticsearch Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants anywhere in the world.
• Take ownership of the entire lifecycle of our search and analytics data platform, including capacity planning, cluster architecture, performance optimization, incident management, migration strategy, and operational excellence.
• Serve as the sole expert on all Elasticsearch and OpenSearch clusters at Chess.com.
• Engage hands-on with cluster internals, create ILM/ISM policies, implement infrastructure changes via GitOps, and make immediate decisions regarding replica allocation when a cluster experiences issues.
• Hold on-call responsibilities for Elasticsearch-related incidents, including cluster health deterioration, node failures, disk pressure, shard imbalance, and write rejection cascades.
• Conduct real-time cluster triage and coordinate with cross-functional teams during production incidents.
• Author post-mortems and implement systemic reliability enhancements.
• Manage snapshots and disaster recovery processes across clusters.
• Provide guidance to engineering teams on index design, mapping strategies, retention policies, and query optimization.
• Oversee access and configuration of Kibana and OpenSearch Dashboards for internal users.
• Define and uphold workload priority tiers across clusters.
• 7+ years of experience operating Elasticsearch at scale, including multi-TB clusters, numerous nodes, and high write throughput.
• In-depth knowledge of Elasticsearch internals, such as segment merging, translog, shard allocation, and managing cluster state.
• Proven production experience with ECK (Elastic Cloud on Kubernetes) or similar operator-based deployments.
• Expertise in Kubernetes operations for stateful workloads, including StatefulSets, persistent storage, and resource management.
• Practical experience in Linux systems administration, focusing on storage and I/O performance.
• Experience managing both Elasticsearch and OpenSearch in production environments, along with informed perspectives on their respective advantages and disadvantages.
• Incident command experience, demonstrating the ability to diagnose and resolve cluster emergencies under pressure while effectively communicating with stakeholders.
• Proficient in Git-based infrastructure management (GitOps), including Helm charts, ArgoCD/Flux, and infrastructure-as-code for cluster configurations.
• Proficient with the Elastic stack APIs covering cluster administration, index templates, data streams, ILM policies, and snapshot/restore functionalities.
• Complete autonomy.
• True scalability.
• Bare metal environments.
• Strategic impact.
• Small team with a high level of trust.
Anduril Industries
Sargent & Lundy
Sargent & Lundy
Get handpicked remote jobs straight to your inbox weekly.