
Senior Platform Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Investigate, troubleshoot, and resolve reliability challenges within the MANTL platform application code, which includes addressing microservice communication failures and correctness issues during failure conditions.
• Tackle failure modes across both third-party and internal system integrations, while developing integration-specific resilience strategies.
• Design, configure, and maintain monitoring systems, dashboards, and alerting mechanisms.
• Implement and enhance distributed tracing across microservices.
• Diagnose and rectify application performance issues, focusing on caching, inefficient code paths, and query performance.
• Strengthen platform and application components through fault injection and resilience testing.
• Build, maintain, and troubleshoot CI/CD build pipelines using GitHub Actions.
• Deploy and resolve issues related to Kubernetes workloads.
• Maintain a roadmap of identified reliability risks that is independent of feature delivery timelines.
• Create documentation and runbooks that detail reliability issues, root causes, and remediation strategies.
• Serve as an escalation point for complex platform reliability challenges.
• Contribute to the definition of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for key platform services.
• 4 to 7 years of experience in software engineering, platform engineering, or a blended role involving development and reliability engineering.
• A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
• Strong expertise in TypeScript with hands-on experience in a Node.js/TypeScript service environment.
• Practical experience in deploying and troubleshooting containerized workloads within Kubernetes.
• Experience in building and maintaining CI/CD pipelines, with a preference for GitHub Actions.
• Proficiency in configuring monitoring systems, dashboards, and alerting within an APM/observability tool, with a preference for Datadog.
• Background in troubleshooting distributed systems and microservice communication failures across various integration points and protocols.
• Familiarity with distributed tracing tools and practices.
• Working knowledge of relational databases and tuning for slow or problematic queries.
• Experience in diagnosing and resolving application performance challenges.
• Skills in designing for and testing failure modes, including fault injection and resilience/chaos-style testing.
• Experience implementing defensive patterns such as idempotency, retries, and circuit breakers.
• Strong analytical and troubleshooting abilities.
• Capability to work independently on ambiguous and reactive reliability issues.
• Ability to clearly communicate technical root causes and remediation plans to both engineering and non-technical stakeholders.
• Experience with message brokers or event-streaming platforms like Kafka is preferred.
• Familiarity with OpenTelemetry or similar distributed tracing frameworks is preferred.
• Previous experience in a regulated or compliance-driven environment is preferred.
• Familiarity with infrastructure-as-code tools such as Terraform is preferred.
• Experience collaborating with infrastructure/SRE teams is preferred.
• Must be eligible to work in the US for full-time employment; employment sponsorship is not available.
• Remote-first work environment.
• Unlimited paid time off.
• 401(k) plan with employer matching.
• A diverse and inclusive workplace.
• A fun organizational culture.
Quantiphi
Agility Technologies Inc
American College of Education
First Due
Get handpicked remote jobs straight to your inbox weekly.