
Senior Software Engineer – IP&R Reliability Engineering
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Design, develop, and sustain automated monitoring solutions, enterprise workload orchestrations, and proactive reliability systems.
• Reduce operational toil by utilizing code.
• Transition legacy systems to scalable, script-based platforms.
• Create centralized health dashboards and set onboarding standards in collaboration with application development teams.
• Optimize alerting pipelines by incorporating PagerDuty, Grafana, custom APIs, and modern observability and incident management tools.
• Oversee and enhance enterprise batch execution flows across CA7, Maestro/HCL TWS, and Rundeck.
• Update and phase out redundant legacy monitoring platforms in favor of script-driven automation.
• Develop, test, deploy, and maintain software along with functional and destructive test suites.
• Work collaboratively with team members and Product Teams within agile frameworks.
• Provide mentorship to junior engineers and contractors in automation scripting, workflow scheduling, troubleshooting, and contemporary software development frameworks.
• Facilitate technical discussions and initiatives.
• Candidates must be eighteen years of age or older.
• Must possess legal authorization to work in the United States.
• Minimum educational requirement: completion of a bachelor's degree program or an equivalent degree in a relevant field.
• At least 3 years of professional work experience is required.
• Preferred: 4+ years of experience in professional software engineering, SRE, DevOps, or system automation within an enterprise setting.
• Strong expertise in scripting and programming languages such as Python, Bash/Shell, Java, Go, or Groovy.
• Experience in developing API integrations and automating operational health dashboards and CLI utilities.
• Practical knowledge of enterprise workload automation and job schedulers like Rundeck, CA7, Maestro/HCL Workload Automation, AutoSys, or Control-M.
• Demonstrated experience in migrating, re-architecting, or modernizing batch job flows and dependency chains.
• Extensive experience configuring alerting, escalation policies, and incident response integrations, including PagerDuty and Slack/Teams webhooks.
• Familiarity with telemetry, log aggregation, and dashboarding tools such as Splunk, Prometheus, Grafana, or Datadog.
• Understanding of retail supply chain operations, warehouse management, or Distribution Center operational systems.
• Experience using AI coding assistants, agentic workflows, or AIOps tools.
• Familiarity with Google Cloud Platform or contemporary containerization platforms such as Docker and Kubernetes.
• Remote/Virtual work arrangement.
• No travel required.
OnePay
Arista Networks
Octus
Tandem Diabetes Care
Get handpicked remote jobs straight to your inbox weekly.