
Staff Platform Engineer – MANTL
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Lead the investigation, troubleshooting, and resolution of intricate reliability challenges within the MANTL platform application code.
• Set the strategic direction for identifying and addressing failure modes across both third-party and internal integrations.
• Take ownership of the design and ongoing development of monitoring, dashboards, and alerting strategies for the platform.
• Establish standards and spearhead the implementation of distributed tracing across microservices.
• Diagnose and resolve application performance issues related to caching, inefficient code pathways, and query performance.
• Define the strategy for fault-injection, resilience testing, and defensive design methodologies.
• Own and enhance the architecture of the GitHub Actions CI/CD build pipeline.
• Guide the deployment and troubleshooting of Kubernetes workloads.
• Define, prioritize, and independently execute a roadmap addressing known reliability risks without being influenced by feature timelines.
• Establish standards for documentation and runbooks regarding reliability issues, root causes, and remediation efforts.
• Act as a senior escalation point for complex platform reliability concerns across various engineering domains.
• Define Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for key platform services and provide insights to leadership on reliability risks and trade-offs.
• Offer technical mentorship and guidance to Platform Engineers.
• 7 to 10 years of experience in software engineering, platform engineering, or a combined development/reliability engineering role.
• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
• Extensive proficiency in TypeScript along with significant production experience in a Node.js/TypeScript service environment.
• Strong hands-on experience in deploying, operating, and troubleshooting containerized workloads in Kubernetes.
• Experience in designing, building, and maintaining CI/CD pipelines; GitHub Actions is preferred.
• Background in designing monitoring, dashboard, and alerting strategies using an APM/observability tool; Datadog is preferred.
• Deep experience in troubleshooting distributed-system and microservice communication failures across various integration points and protocols.
• Strong expertise with distributed tracing tools and practices at scale.
• Familiarity with relational databases for diagnosing complex query and schema-level performance issues.
• Proven ability to resolve application performance issues related to caching, inefficient code pathways, and slow queries.
• Experience in designing and leading failure-mode testing, fault injection, and resilience/chaos-style testing.
• Experience in implementing defensive patterns such as idempotency, retries, and circuit breakers.
• Capability to work independently on ambiguous, high-impact reliability challenges and set technical direction.
• Excellent communication skills for articulating technical root causes, risks, and remediation strategies to both technical and non-technical leadership.
• Experience in mentoring fellow engineers.
• Candidates must be eligible to work in the US for full-time employment.
• Preferred: experience with Kafka or event-streaming platforms.
• Preferred: familiarity with OpenTelemetry or similar tracing frameworks.
• Preferred: experience in a regulated or compliance-driven environment.
• Preferred: familiarity with Terraform or similar infrastructure-as-code tools.
• Preferred: experience collaborating with infrastructure/SRE teams.
• Preferred: experience influencing engineering roadmap prioritization.
• Remote-first work environment.
• Unlimited paid time off.
• 401(k) plan with employer matching.
• Commitment to a diverse and inclusive workplace.
• A fun and engaging company culture.
Quantiphi
Agility Technologies Inc
American College of Education
First Due
Get handpicked remote jobs straight to your inbox weekly.