
Senior Sustaining and Forward Deployed Engineer
Posted Jul 7

Posted Jul 7
This is a fully remote position, open to applicants in United States.
• Serve as a senior technical escalation point during production incidents.
• Lead real-time incident triage, mitigation, and recovery efforts.
• Drive root cause analysis (RCA) with an emphasis on systemic, long-term solutions.
• Identify recurring failure patterns and advocate for architectural or operational enhancements.
• Collaborate with Customer Success and Engineering to manage customer impact during incidents.
• Take ownership of post-launch reliability, stability, and operational quality of core systems.
• Investigate and resolve intricate field issues and production defects.
• Ensure that fixes developed during incidents or customer escalations are integrated into the core product.
• Enhance operational readiness of services through improved runbooks, monitoring, and alerting.
• Minimize operational toil by automating repetitive manual tasks.
• Engage directly with strategic customers to address real-world, production-grade technical challenges.
• Support complex deployments, integrations, and escalations within customer environments.
• Act as a trusted technical partner to customers during high-impact situations.
• Translate customer insights into tangible product, platform, and operational enhancements.
• Contribute to reusable tools, playbooks, and best practices that facilitate future deployments.
• Serve as a subject matter expert for AWS-hosted production systems.
• Troubleshoot and resolve issues across:
• AWS compute, storage, networking, IAM, and security.
• Databricks jobs, clusters, and Spark-based data pipelines.
• Debug performance degradation, scalability challenges, job failures, and data accuracy issues.
• Collaborate with platform and data teams to strengthen systems for reliability, scalability, and operability.
• Write production-quality code to:
• Automate operational workflows.
• Enhance reliability and observability.
• Eliminate manual intervention and decrease incident frequency.
• Primarily contribute in Python, with some exposure to JVM-based systems as required.
• Review code with a strong focus on operability, resiliency, and maintainability.
• Promote “build it so it can be operated” engineering standards.
• Provide technical leadership without formal authority, influencing design and operational decisions.
• Mentor engineers through pairing, reviews, and incident leadership.
• Work closely with Product, Engineering, Data, and Customer teams.
• Operate effectively in high-pressure, ambiguous environments, particularly during customer-impacting incidents.
• Over 10 years of experience in software engineering, SRE, sustaining engineering, or production operations.
• Extensive hands-on experience operating production systems in AWS.
• Strong expertise in troubleshooting Databricks and large-scale data platforms.
• Proficient in Python with experience in building production services or tools.
• Solid understanding of:
• Distributed systems.
• Incident management and RCA practices.
• Monitoring, alerting, and observability.
• CI/CD pipelines that leverage Infrastructure as Code.
• Proven track record of owning problems end-to-end, from detection to permanent resolution.
• Excellent communication skills, particularly during incidents and customer escalations.
• Ability to work backward from customer impact to root cause across systems and codebases, delivering fixes in environments with minimal documentation.
• Strong instinct for operational risk, with the capability to proactively identify failure modes and strengthen systems before they affect customers.
• Unlimited paid time off – recharge when you need it.
• Work from anywhere – flexibility to fit your life.
• Comprehensive health coverage – multiple plan options to choose from.
• Equity for every employee – share in our success.
• Growth-focused environment – your development matters here.
• Home office setup allowance – one-time support to get you started.
• Monthly cell phone allowance – stay connected with ease.
Future Connections
Akamai Technologies
cigus
Digi International
Get handpicked remote jobs straight to your inbox weekly.