DevOps Engineer IV – Infrastructure

Posted 3 days ago

This is a fully remote position, open to applicants in India.

📋 Description

• Take charge of the cloud infrastructure and DevOps operations for Astreya's product offerings.

• Design, construct, and manage the GCP environment, encompassing GKE, Artifact Registry, Cloud SQL, Cloud Storage, Memorystore, Vertex AI, monitoring/logging, and DNS.

• Oversee network architecture, perimeter security, secrets management, key management, identity, and certificate lifecycle design.

• Assess new platform services and generate recommendations supported by functional proofs of concept.

• Manage cloud cost allocation, budget notifications, rightsizing, and planning for commitments.

• Ensure that product deployment into customer-VPC and Astreya-hosted environments is repeatable, well-documented, and dependable.

• Supervise Helm charts, namespace and pod topology, environment configurations, migrations, and rollback strategies.

• Collaborate with customer infrastructure and security teams during onboarding and facilitate the transition of products to production.

• Create and maintain architecture, data flow, access control, encryption, and audit documentation.

• Manage CI/CD processes across products, including build pipelines, artifact promotion, environment strategies, and branching methodologies.

• Develop Terraform automation and internal tools for infrastructure provisioning and operations.

• Establish observability through metrics, logs, traces, alerting, dashboards, and SLOs.

• Enhance platform scalability, reliability, capacity, and performance, including infrastructure for AI/LLM.

• Strengthen platform security through image scanning, vulnerability management, network policies, secrets hygiene, and patching.

• Lead and mentor DevOps and infrastructure engineers, establish standards, and review their work.

• Collaborate with product owners and engineering leads to identify and resolve infrastructure bottlenecks.

• Manage platform incident responses and conduct thorough root cause analyses.

• Refine support processes and tools, document decisions, and maintain runbooks and resolution records.


⛳️ Requirements

• Bachelor’s degree (B.S/B.A) from a recognized four-year institution and more than 8 years of relevant experience and/or training; or an equivalent blend of education and experience.

• Extensive hands-on experience with Google Cloud, particularly in designing and operating production workloads on GCP.

• Strong experience with production Kubernetes, preferably GKE, including networking, ingress, autoscaling, resource management, and troubleshooting failures under load.

• Professional experience with infrastructure as code using Terraform, demonstrating module design, state management, and review practices.

• Proficient scripting and coding skills in Python, Go, Bash, or similar languages.

• Sufficient software development experience to navigate application repositories effectively.

• Solid expertise in cloud networking and security, including VPC design, private connectivity, IAM, secrets management, and TLS.

• Proven ownership of CI/CD processes across multiple products and release cycles, with source control branching strategies.

• Experience with monitoring, alerting, and incident management, including the creation of RCAs.

• Familiarity with open-source tools within large distributed systems.

• Strong communication skills for collaborating with customer security teams and developers.

• Ability to swiftly learn new technologies and manage multiple tasks simultaneously.

• Basic knowledge of Azure; experience with AWS is advantageous.

• Experience deploying products in customer-controlled cloud environments from a vendor perspective.

• Experience supporting AI/ML workloads in a production setting.

• Familiarity with SOC 2, ISO 27001, or similar infrastructure audits.

• Experience integrating with enterprise ITSM platforms such as ServiceNow.

• Previous experience as a first or early infrastructure hire on a product team.

• Relevant certifications such as Google Professional Cloud Architect, Professional Cloud DevOps Engineer, or CKA.

• Capability to perform office-related tasks, including prolonged sitting or standing.

• Ability to navigate within an office environment.

• Proficient in using a computer.

• Effective communication skills.

• Ability to perform occasional repetitive wrist, hand, or finger movements as needed.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work opportunities.

• Professional development and continuous learning opportunities.

• Generous vacation and paid time off policy.

People also viewed

FourEnergy GmbH13 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF15 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam19 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers