
Senior Platform Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Design, provision, and maintain cloud infrastructure utilizing Infrastructure as Code, while managing development, staging, and production environments with repeatable and auditable configurations.
• Take ownership of and enhance CI/CD pipelines featuring automated testing gates, blue/green and canary rollouts, along with dependable rollbacks.
• Build and manage observability across metrics, logging, distributed tracing, dashboards, and alerting; establish SLOs, SLIs, and error budgets.
• Lead incident response efforts, which include on-call participation, triage, mitigation, communication, and conducting blameless postmortems.
• Automate operational challenges through scripting and tools to enhance detection and recovery times.
• Design systems for scalability and resilience through capacity planning, load and failure testing, autoscaling, redundancy, disaster recovery, and backup strategies.
• Manage Docker and Kubernetes workloads, encompassing networking, service discovery, and resource management.
• Collaborate with Security and Engineering teams to strengthen infrastructure through secrets management, network segmentation, encryption, vulnerability scanning, patching, and audit logging in accordance with HIPAA-aligned PHI requirements.
• Work alongside development teams to incorporate reliability and operability into services from the design phase through to production.
• A Bachelor’s degree in computer science, software engineering, or a related field, or equivalent practical experience.
• Over 5 years of experience in Site Reliability, DevOps, or infrastructure engineering roles within startup or growth-stage companies, preferably in healthcare, health tech, or other regulated, high-availability, data-sensitive industries.
• Extensive hands-on experience managing production systems on a leading cloud platform (AWS, GCP, or Azure), including compute, networking, storage, and managed database services.
• Strong expertise in Infrastructure as Code (Terraform or similar) and configuration management.
• Experience in containerization and orchestration (Docker and Kubernetes), including deployment, scaling, and troubleshooting.
• Proven ability in building CI/CD pipelines and release automation, with a history of facilitating safe and frequent deployments.
• Practical experience with observability and monitoring tools (e.g., Datadog, Prometheus, Grafana, CloudWatch, ELK/OpenSearch) and in defining SLOs, SLIs, and error budgets.
• Strong scripting and automation capabilities (Python, Go, or Bash) and familiarity with command-line and cloud CLIs.
• Experience in leading incident response and on-call duties, including postmortem and root-cause analysis methodologies.
• Solid knowledge of infrastructure and network security, secrets management, and encryption, with experience in handling sensitive or protected data (PHI/HIPAA experience is highly preferred).
• Proficient in Git and version control workflows; familiarity with the Agile Development Framework and ideally the Atlassian toolset (Jira and Confluence).
• Excellent verbal and written communication abilities; calm, methodical, and detail-oriented under pressure, and comfortable navigating ambiguity in an early-stage setting.
• Significant ownership through stock options, enabling you to directly benefit from the value you help create as we expand.
• Competitive base salary.
• Comprehensive medical, dental, and vision insurance options, with contributions from Rezilient towards your premiums.
• Free access to Rezilient's clinical programs for you and your household members, offering the same connected care provided to our patients.
• A 401(k) retirement plan to assist in your long-term financial security.
• Flexible Paid Time Off, allowing you to recharge without monitoring days.
• Dedicated Paid Sick Leave, distinct from your FTO, prioritizing health needs.
• 11 paid company holidays annually.
• Paid family leave to support you during significant life events.
• Optional ancillary benefits, including life insurance, disability coverage, and a Health Savings Account (HSA).
Board Intelligence
T-Rex Solutions, LLC
Teladoc Health
IMH
Get handpicked remote jobs straight to your inbox weekly.