Senior/Staff DevOps Engineer

Posted Sep 14

This is a fully remote position, open to applicants in Portugal.

📋 Description

• Establish and drive the technical vision along with a quarterly roadmap for infrastructure, ensuring clear trade-offs and measurable objectives.

• Manage and enhance AWS and Kubernetes (EKS) infrastructure, encompassing cluster management, autoscaling with Karpenter, policy enforcement with Kyverno, and operations without downtime.

• Take ownership of Infrastructure as Code from start to finish, utilizing Terraform and AWS CDK in TypeScript.

• Oversee GitLab CI/CD, which includes reusable/shared templates, OIDC, and self-managed GitLab.

• Develop and maintain practical observability solutions using Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log pipelines, and APM.

• Ensure rapid blue-green deployments and health-gated automated rollback processes.

• Manage zero-downtime PostgreSQL schema migrations through expand/contract and CI migration gating.

• Lead security engineering efforts in a HIPAA-compliant environment, focusing on secrets hygiene, credential rotation, short-lived credentials, leak scanning, and PHI-aware log and data management, with Vault managed as code.

• Collaborate directly with product teams to eliminate infrastructure friction and enhance the developer experience.

• Incorporate agentic AI as an integral part of the workflow and integrate autonomous-agent output into production environments.

• Define technical direction, execute infrastructure projects hands-on, and take responsibility for reliability, cost, security, performance, deployment health, and developer experience outcomes.


⛳️ Requirements

• Minimum of 6 years experience in DevOps or infrastructure engineering.

• Strong foundational knowledge of systems.

• Proficient in Linux administration and troubleshooting, including performance analysis, resource management, and process debugging.

• Hands-on experience with AWS services such as EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, and S3.

• Proven experience with production Kubernetes/EKS, including cluster management, node scaling, and policy enforcement; familiarity with Karpenter, Kyverno, or similar tools.

• Strong experience in Infrastructure as Code using Terraform and AWS CDK in TypeScript.

• Ownership of CI/CD processes with GitLab CI/CD, including reusable/shared templates, OIDC id_tokens, and self-managed GitLab.

• Practical knowledge of monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log shipping, and error tracking/APM.

• Hands-on experience in security engineering, focusing on secrets rotation, short-lived credentials, leak scanning, and PHI-aware logging.

• Experience with HashiCorp Vault as code, including KV and JWT/OIDC authentication for CI and policy design.

• Familiarity with blue-green deployments featuring automated, health-gated rollback.

• Experience with PostgreSQL zero-downtime schema migrations utilizing expand/contract and migration gating in CI.

• Knowledge of container technologies, specifically Docker, ECR, immutable tags, and image lifecycle management.

• Understanding of network and protocol fundamentals, including load balancing, TLS, and DNS.

• Practical experience with agentic AI workflows, such as Claude Code or similar technologies.

• Developer-centric mindset with strong problem-solving abilities for complex system challenges.

• Excellent technical writing skills, including documentation-as-code, ADRs, and design documents via merge requests.

• Proficiency in both Russian and English (B1 level).

• Experience working effectively within remote, distributed teams.

• Preferred: experience in regulated or compliance-heavy environments such as HIPAA or SOC 2.

• Preferred: familiarity with Ansible for VM fleet management.

• Preferred: knowledge of Node.js application operations using pm2 and npm.

• Preferred: experience with GitOps tooling like ArgoCD or Flux and deeper PostgreSQL database management capabilities.

• Preferred: AWS certifications.


🏝️ Benefits

• Competitive compensation package.

• Health insurance provided after the probation period.

• Compensation for sports and wellness activities.

• Personalized English lessons offered through Preply.

• 19 paid vacation days granted annually.

• 4 additional wellness days each year.

• Paid sick leave for the first 5 working days.

• Thoughtful gifts for significant life events.

• Offline corporate events.

• Fully remote long-term collaboration under a B2B model.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers