Site Reliability Engineer III

atVida HealthRemoteUS flagUnited StatesFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$175k – $185k/year

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Integrate Terraform across both infrastructure and data repositories into a streamlined, well-documented format.

• Set conventions for state management, module organization, code reviews, and CI checks.

• Standardize environments and enhance build and deployment automation using GitHub Actions.

• Implement drift detection and notification mechanisms.

• Execute patches and upgrades for Cloud SQL databases and application runtimes.

• Optimize compute and database workloads for future growth, including connection pooling and scaling enhancements.

• Assess Kubernetes architecture, determining the right timing for transitioning to a multi-cluster setup.

• Enhance monitoring and observability capabilities in Datadog and Cloud Monitoring.

• Design access for contractors and external partners while ensuring the protection of health information.

• Phase out outdated infrastructure and tools.

• Establish operational processes, including runbooks, on-call rotations, and escalation documentation.

• Ensure infrastructure readiness for Vida's enterprise launches on January 1.

• Modernize, consolidate, and scale infrastructure.

• Join the Enablement Team, report to the Engineering Manager, and collaborate closely with the Lead Engineer.

• Contribute to the development of SRE practices at Vida.


⛳️ Requirements

• A minimum of a Bachelor's degree.

• Over 5 years of experience in SRE, DevOps, or infrastructure engineering with significant ownership of production systems.

• Extensive hands-on experience with Terraform, including module structuring and state management across environments.

• Strong knowledge of GCP, encompassing GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking, load balancing, and cost management.

• Experience with production Kubernetes, including autoscaling, resource management, and decision-making regarding cluster resources.

• Practical experience in building monitoring, alerting, and dashboards using tools such as Datadog or Cloud Monitoring.

• Proficient in Python for development and automation tasks.

• Capable of collaborating across various teams and fields, effectively communicating infrastructure decisions to non-specialists.

• Applicants must be authorized to work in the U.S.

• Vida Employees must be located in or able to work from the U.S.; international work is not permitted.

• Must reside in or be able to work from a state where Vida is registered.

• Preferred: Experience as an early or first SRE hire.

• Preferred: Experience in refactoring or consolidating a large, organically developed Terraform codebase.

• Preferred: Experience enhancing observability from a less developed baseline.

• Preferred: Experience in a HIPAA-regulated or other compliance-focused environment.

• Preferred: CI/CD experience using GitHub Actions.

• Preferred: Experience deploying Django applications or Airflow in production on Kubernetes.

• Preferred: Experience designing or migrating to multi-cluster Kubernetes architectures.


🏝️ Benefits

• Fully remote work with no restrictions related to time zones.

• An Equal Employment Opportunity and Affirmative Action employer.

• Reasonable accommodations available for qualified individuals with disabilities and disabled veterans.

• No work visa sponsorship is required or available.

People also viewed

Capgemini1 day ago

Senior DevOps Engineer – Linux, Python, Bash

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Satellite Office1 day ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RELX1 day ago

Senior Site Reliability Engineer

US flagMassachusetts, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)$95.3k – $158.8k/year
ApplyView job
KnowBe41 day ago

Staff Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
KnowBe41 day ago

Senior Site Reliability Engineer, Remote in Brazil

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
KnowBe41 day ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $155k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers