
Director β Platform Engineering
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in Maryland.
β’ Take full ownership of the complete architecture of the Elastic estate, which includes deployment models, cluster topology, tiering, and multi-tenant isolation.
β’ Have final authority on decisions related to platform architecture.
β’ Lead the transition of the remaining ECE-based estate to ECK.
β’ Implement the existing engineering plan, address technical gaps, and ensure the program is completed without service interruptions.
β’ Stay actively involved in design reviews, upgrade planning, escalated troubleshooting, and prototype development.
β’ Oversee and refine capacity planning; predict ingest growth, size clusters and hardware, and drive procurement processes.
β’ Lead the weekly Platform Engineering change control meetings and represent the department at the enterprise Change Advisory Board.
β’ Manage ingest pipeline design and health, log source integration, and data quality benchmarks.
β’ Collaborate with Detection Engineering and Security Operations on enhancing detection and alerting quality.
β’ Plan and implement Elastic version upgrades with minimal disruption to clients.
β’ Maintain platform currency and manage risks associated with end-of-life components.
β’ Act as the designated Platform Engineering contact during Major Incident Management.
β’ Direct root cause analysis and resolution for incidents originating from the platform.
β’ Guide, mentor, and develop platform engineers; manage hiring, performance evaluations, career growth, and workload distribution.
β’ Oversee the operational cadence of the department.
β’ Lead the Engineering Leadership Roundtable across engineering, infrastructure, detection engineering, security operations, and product teams.
β’ Communicate platform constraints to product and commercial stakeholders.
β’ Manage relationships with Elastic, including licensing, renewals, true-ups, tool evaluations, and technical negotiations.
β’ Drive the integration of acquired environments, covering aspects such as cluster migration, ingest re-pointing, client mapping, and license consolidation.
β’ Establish standards for documentation regarding architecture, runbooks, and recovery processes.
β’ Proactively reduce reliance on key individuals as a specific early deliverable.
β’ Act as the accountable officer for platform-related control documentation and support SOC 2 and client audit evidence requests.
β’ At least 10 years of experience in infrastructure, platform, or systems engineering.
β’ Minimum of 4 years in leadership roles within engineering teams.
β’ Proven track record of ownership over a large-scale, multi-tenant data platform in a production environment.
β’ Experience working within a managed services, MSSP, or MDR setting where platform availability is contractually obligated.
β’ History of inheriting and stabilizing environments created by others, including reverse-engineering undocumented systems.
β’ Successful record of delivering a platform migration or consolidation program from start to finish.
β’ Extensive hands-on experience with the Elastic Stack at scale β including Elasticsearch, Kibana, Logstash, Beats, and Agent β particularly in multi-tenant cluster design and operation.
β’ Strong preference for experience with Elastic Cloud on Kubernetes (ECK).
β’ Solid understanding of log ingestion architecture, data normalization (ECS), index lifecycle management, and retention strategies.
β’ Familiarity with SIEM and detection engineering concepts; knowledge of detection-as-code methodologies is a plus.
β’ Proficient in Linux systems administration.
β’ Experience with containerization and orchestration, particularly Kubernetes.
β’ Knowledge of infrastructure-as-code practices.
β’ Expertise in identity and access integration: SAML, LDAP, SSO, and role-based access design.
β’ Skills in capacity modeling and cost-to-serve analysis for data-intensive platforms.
β’ Ability to maintain architectural authority while remaining receptive to alternate viewpoints.
β’ Capability to lead the execution of plans developed by others and build team trust.
β’ Comfortable managing a department where most tasks are unplanned while safeguarding strategic capacity.
β’ Clear communication skills, both written and verbal, with engineers and executives.
β’ Effective in a distributed organization across multiple time zones.
β’ Elastic certification is preferred, such as Elastic Certified Engineer or Elastic Certified Architect.
β’ Experience in 24x7 operational settings and formal change management frameworks like ITIL is preferred.
β’ Prior experience presenting platform status to enterprise clients or in audit situations is preferred.
β’ Flexible Paid Time Off.
β’ 401k plan with a company match.
β’ Comprehensive Medical, Dental, and Vision Coverage.
β’ Voluntary Short Term and Long-Term Disability options.
β’ Employee Assistance Program that includes a Mental Health Supplement.
β’ Voluntary Basic, Accidental, and other ancillary life insurance options.
β’ Contribution to Health Savings Account (with selection of a High Deductible Health Plan).
β’ 10 annual paid holidays.
β’ Fully remote work arrangement.
OPENDataJobs
Presidio
EasyLlama - HR & Compliance Training For Modern Teams
Coinbase
Get handpicked remote jobs straight to your inbox weekly.