
Site Reliability Engineer III
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Integrate Terraform across both infrastructure and data repositories into a streamlined, well-documented format.
• Set conventions for state management, module organization, code reviews, and CI checks.
• Standardize environments and enhance build and deployment automation using GitHub Actions.
• Implement drift detection and notification mechanisms.
• Execute patches and upgrades for Cloud SQL databases and application runtimes.
• Optimize compute and database workloads for future growth, including connection pooling and scaling enhancements.
• Assess Kubernetes architecture, determining the right timing for transitioning to a multi-cluster setup.
• Enhance monitoring and observability capabilities in Datadog and Cloud Monitoring.
• Design access for contractors and external partners while ensuring the protection of health information.
• Phase out outdated infrastructure and tools.
• Establish operational processes, including runbooks, on-call rotations, and escalation documentation.
• Ensure infrastructure readiness for Vida's enterprise launches on January 1.
• Modernize, consolidate, and scale infrastructure.
• Join the Enablement Team, report to the Engineering Manager, and collaborate closely with the Lead Engineer.
• Contribute to the development of SRE practices at Vida.
• A minimum of a Bachelor's degree.
• Over 5 years of experience in SRE, DevOps, or infrastructure engineering with significant ownership of production systems.
• Extensive hands-on experience with Terraform, including module structuring and state management across environments.
• Strong knowledge of GCP, encompassing GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking, load balancing, and cost management.
• Experience with production Kubernetes, including autoscaling, resource management, and decision-making regarding cluster resources.
• Practical experience in building monitoring, alerting, and dashboards using tools such as Datadog or Cloud Monitoring.
• Proficient in Python for development and automation tasks.
• Capable of collaborating across various teams and fields, effectively communicating infrastructure decisions to non-specialists.
• Applicants must be authorized to work in the U.S.
• Vida Employees must be located in or able to work from the U.S.; international work is not permitted.
• Must reside in or be able to work from a state where Vida is registered.
• Preferred: Experience as an early or first SRE hire.
• Preferred: Experience in refactoring or consolidating a large, organically developed Terraform codebase.
• Preferred: Experience enhancing observability from a less developed baseline.
• Preferred: Experience in a HIPAA-regulated or other compliance-focused environment.
• Preferred: CI/CD experience using GitHub Actions.
• Preferred: Experience deploying Django applications or Airflow in production on Kubernetes.
• Preferred: Experience designing or migrating to multi-cluster Kubernetes architectures.
• Fully remote work with no restrictions related to time zones.
• An Equal Employment Opportunity and Affirmative Action employer.
• Reasonable accommodations available for qualified individuals with disabilities and disabled veterans.
• No work visa sponsorship is required or available.
Capgemini
Satellite Office
RELX
KnowBe4
Get handpicked remote jobs straight to your inbox weekly.