
Platform Engineer – Incident Management
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Take charge of on-call shifts from start to finish: promptly acknowledge alerts within the Service Level Agreement (SLA), declare incidents, assign roles, maintain a communication rhythm, and guide incidents to successful resolution.
• Develop automation for the Sev1/Sev2 postmortem process, which includes scheduling, sending facilitation reminders, assigning action items, tracking ownership, and enforcing due dates.
• Utilize AI to detect patterns in incidents and recommend systemic improvements, such as enhancements to runbooks, alert fine-tuning, platform fortification, and modifications to processes.
• Create and enhance internal automation and tools, including AI-driven incident response workflows, aimed at minimizing manual labor and speeding up detection and resolution.
• Contribute to and enhance observability through dashboards, alert configurations, and early-warning signals across Bitso’s platform.
• Engage in change and maintenance management processes, applying risk management strategies to mitigate deployment-related incidents.
• Collaborate with engineering teams throughout the organization to identify platform risks and implement preventive measures.
• Ensure that incident management tools, runbooks, and severity criteria remain accurate, up-to-date, and beneficial for the wider engineering community.
• Report directly to the Incident Management Manager.
• Demonstrated ability to operate effectively in high-pressure incident situations, including clear communication with senior stakeholders and leadership while production issues are active.
• Practical experience with Kubernetes — comfortable with deploying, debugging, and addressing pod-level challenges.
• Strong understanding of CI/CD pipelines and contemporary DevOps methodologies.
• Background in software development in any programming language; the capability to read, write, and debug code is essential (experience in Python or Java is advantageous).
• A robust automation mindset: capable of identifying repetitive tasks and eliminating them.
• Experience in building or working with AI agents or workflows based on Large Language Models (LLMs) is highly desirable.
• Exceptional interpersonal and written communication skills.
• A self-motivated learner who can contribute without needing a fully defined path.
• Experience in the fintech or cryptocurrency sectors is a plus.
• Full ownership of the incident lifecycle and incident management experience is implied by the responsibilities of this role.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and continuous learning.
• Flexible work hours and remote work options.
• Collaborative and innovative work environment.
Helpware
Capgemini
Peraton
Phantom
Get handpicked remote jobs straight to your inbox weekly.