
Production Support Engineer
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in Argentina, +2 more countries.
• Oversee the health of the platform while handling incoming incidents.
• Differentiate between critical issues like outages and service degradation versus non-critical bugs and defects.
• Analyze incidents utilizing logging tools, API calls, responses, and log information.
• Take ownership of the incident management response process from the initial alert through to post-mortem evaluations and follow-up on corrective actions.
• Proactively inform stakeholders about critical issues, risks to service level agreements (SLAs), and updates via email, phone, or ticketing systems.
• Cross-check tickets across various systems and track defects until resolution.
• Convey technical findings in accessible language to partners and call center teams.
• Manage the incident queue in Jira, prioritizing bugs during engineering sprint cycles.
• Engage in weekly cross-functional meetings with engineering and account/call center management.
• Propose ongoing enhancements to applications and processes.
• Participate in on-call rotations following ramp-up and respond to OpsGenie alerts within specified SLA timeframes.
• Ensure the health and reliability of a revenue-critical sales platform that connects end users with engineering.
• Minimum of 2 years of experience in troubleshooting and resolving issues related to applications, servers, or infrastructure environments.
• At least 2 years of experience in providing clear status updates regarding tasks, issues, and resolutions to stakeholders at various levels.
• Proficient understanding of APIs, capable of reading and interpreting API calls and responses.
• Familiarity with Postman or similar API testing tools.
• Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic.
• Experience with SQL queries for troubleshooting and generating ad hoc reports.
• Basic proficiency in reading HTML and JSON as well as utilizing browser developer tools.
• Capable of participating in technical bridge calls and following incidents to resolution.
• Excellent communication skills suitable for both technical and non-technical audiences.
• Strong time management, prioritization, and organizational abilities under pressure.
• Customer service-oriented mindset with a genuine interest in assisting end users.
• Traits of empathy, humility, and comfort with uncertainty.
• Availability during U.S. Eastern business hours (9 AM–6 PM ET); candidates in the Eastern timezone are strongly preferred for onboarding and on-call coordination.
• Bachelor's degree in a relevant field or equivalent work experience.
• Preferred: AWS Cloud Practitioner certification or higher.
• Preferred: Familiarity with Git.
• Preferred: Knowledge of Terraform or similar infrastructure-as-code concepts.
• Preferred: Experience with OpsGenie or comparable alerting platforms.
• Preferred: AI tools for troubleshooting and investigation workflows.
• Preferred: Understanding of the engineering deployment lifecycle and release processes.
• Preferred: Knowledge of on-call rotation structures and incident severity frameworks.
• Preferred: Experience in managing communication with partners or call centers during live incidents.
• Preferred: Background in travel, hospitality, or high-volume transactional platforms.
• Comprehensive onboarding documentation and a structured 6-month ramp to achieve full self-sufficiency.
• On-call rotations commence only when you are prepared, with managerial support during initial rotations.
• Active alert window is from 8 AM to 1 AM Eastern; overnight suppression windows are included.
• Collaborate with a global team across diverse continents and cultures.
• An inclusive environment that emphasizes continuous learning, innovation, and ethical AI standards.
NVIDIA
SAIC
SAIC
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.