
Staff Application Support Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Diagnose and resolve intricate production challenges across CINC's applications, services, integrations, and data layers.
• Serve as a senior escalation point for Level 2 support, delivering in-depth technical investigations and clear resolution strategies.
• Engage in and lead incident response efforts, encompassing triage, mitigation, recovery, and post-incident evaluations.
• Provide calm, organized, and timely communication during incidents to both internal stakeholders and customer-facing teams.
• Ensure comprehensive documentation of incidents, including root cause analysis, impact assessment, and preventive measures.
• Identify repetitive or high-volume support tasks and eliminate them through automation, with scripting being your primary approach.
• Develop and maintain internal tools that enhance investigation speed, streamline case management, and reveal operational insights.
• Create scripts in Python, Bash, or similar languages to automate data retrieval, diagnostics, remediation steps, and reporting workflows.
• Design and manage Datadog monitors, dashboards, alert routing rules, and synthetic checks that provide early alerts and minimize noise.
• Implement and enhance AI-enabled support platforms, such as Intercom Fin, to deflect issues and elevate response quality.
• Treat support automation and tooling as core engineering tasks, ensuring appropriate documentation, version control, and maintainability.
• Develop and uphold knowledge bases that facilitate quicker resolutions and enable customer self-service.
• Collaborate within and help advance CINC's Datadog environment, including monitors, dashboards, alert routing, log pipelines, and observability workflows.
• Design alerting and escalation protocols that highlight issues early and mitigate customer impact.
• Create and maintain high-quality operational documentation, runbooks, and incident response playbooks.
• Partner with platform and engineering teams to enhance observability, diagnostics, and operational preparedness.
• Maintain strong feedback loops between customers, application support, and Level 3 product engineering teams.
• Identify recurring issues, failure trends, and systemic risks, providing clear recommendations for remediation.
• Assist in ensuring that insights from incidents lead to lasting product and process enhancements.
• Advocate for operational excellence and reliability as shared responsibilities among the engineering teams.
• Over 10 years of experience in application support, production support, or support engineering roles.
• Extensive expertise in incident management, operational workflows, and production environments.
• Strong practical experience with Datadog, including monitors, dashboards, log management, and alerting mechanisms.
• Proven scripting skills (Python, Bash, or similar) for developing tools, automating workflows, and optimizing support operations.
• Demonstrated ability to troubleshoot complex SaaS systems spanning application, integration, and data layers.
• Experience in designing or working with alerting and monitoring systems.
• Exceptional written and verbal communication skills, particularly in high-pressure scenarios.
• Ability to work autonomously, prioritize tasks, and function effectively in a fast-paced, distributed environment.
• Great working environment.
• Opportunity for growth.
• Learning & development resources.
• Work/life balance.
• Medical.
• Dental.
• Vision.
• Life Insurance.
• Short-term & Long-term disability insurance.
• Flexible Spending Plan (childcare & healthcare).
• 401K (matching available).
• 136 hours of paid time off per year.
• 10 paid holidays per year.
• 2 self-care/mental health days per year.
• Your birthday is a paid holiday!
• Free Coca-Cola beverages and snacks (when in office).
Avnet
Teradyne
Intetics
New Charter Technologies
Get handpicked remote jobs straight to your inbox weekly.