
Software Engineer – Site Reliability
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Develop and enhance tools, services, and automation to ensure the reliability of production systems.
• Create shared observability functionalities for metrics, logs, traces, service health, and customer impact.
• Enhance incident response and operational preparedness through tooling, automation, standards, and production signals.
• Establish resiliency features to detect failure modes, mitigate operational risks, and recover from infrastructure or application failures.
• Substitute repetitive manual operational tasks with robust software and automation solutions.
• Implement AI in software development and operational problem-solving processes.
• Spot opportunities for AI-driven capabilities that enhance incident investigation, observability, reliability, and engineering efficiency.
• Deliver well-defined engineering projects independently and contribute to technical designs.
• Collaborate with product engineering, Cloud Platform, Delivery, Developer Platform, Security, and infrastructure teams.
• A minimum of 3 years of professional experience in software engineering, site reliability engineering, or a related field.
• Proficient in software development using one or more general-purpose programming languages such as Python, Go, JavaScript, or TypeScript.
• Experience in designing, building, testing, and maintaining production software, internal tools, or infrastructure.
• Familiarity with cloud infrastructure, distributed systems, observability, monitoring, or production operations.
• Experience in on-call duties or incident response for production systems.
• Ability to independently manage well-defined engineering projects, navigate technical ambiguities, and collaborate with various engineering teams.
• Experience utilizing AI-assisted development tools at different stages of the software engineering lifecycle.
• Preferably skilled in Kubernetes, AWS, infrastructure as code, and cloud-native production environments.
• Preferably experienced with observability platforms like Datadog, Sumo Logic, or CloudWatch.
• Preferably knowledgeable about reliability practices including service level objectives, capacity planning, resiliency testing, disaster recovery, or operational readiness.
• Competitive compensation package, inclusive of base salary, bonus opportunities, and annual equity grants that vest quarterly.
• 401(k) or Group Retirement Savings Plan featuring a company match of $2 for each $1 contributed, up to $15,000 annually.
• Employee Stock Purchase Plan (ESPP) offering discounted stock purchase options for eligible employees (US only).
• Comprehensive medical, dental, vision, and wellness resources for US employees.
• Additional health coverage for Canadian employees.
• Contributions to Health Savings Accounts for eligible plans (US only).
• Life insurance and disability coverage.
• Paid time off, sick leave, and company holidays.
• Paid family and parental leave.
• Family-oriented fertility, parenthood, and caregiving benefits.
• Employee Assistance Program providing mental health support and life-centered resources.
• Financial planning tools and concierge service (US only).
• Annual wellness allowance.
• Annual productivity allowance for relevant tools and resources.
• Team events, all-company updates, and employee resource groups.
• Catered lunches and stocked micro-kitchens in offices located in the Bay Area, Austin, Columbus, and New York City.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.