DevOps Engineer

Posted 16 hours ago

This is a fully remote position, open to applicants in Portugal.

📋 Description

• Take ownership of Colonist’s infrastructure from start to finish, encompassing Kubernetes clusters, Docker builds, and DigitalOcean services.

• Manage cluster and cloud resources through code using Terraform.

• Transition production deployments from manual weekly releases to safer, more frequent updates.

• Oversee capacity management, including autoscaling, load balancer thresholds, and connection limits.

• Create and maintain documentation for systems, architectural decisions, runbooks, and incident reports.

• Enhance the workflow from code merge to production deployment.

• Operate PostgreSQL, Redis, and NATS under actual production load.

• Analyze slow queries, as well as connection and capacity constraints, schema migrations, and message compatibility during rolling deployments.

• Manage Cloudflare DNS settings, including proxying, caching, rate limiting, Turnstile, and R2 configurations.

• Ensure observability through Sentry and New Relic while consolidating monitoring tools.

• Develop self-service infrastructure pathways and guardrails.

• Address incidents, lead root cause analyses, and draft preventive postmortems.


⛳️ Requirements

• Senior-level expertise in DevOps.

• Experience with production Kubernetes, including rollouts, autoscaling, networking, load balancers, and troubleshooting problematic deployments.

• Documentation-first approach to systems, decisions, and runbooks.

• Proficient in Linux, Docker, and a CI system not developed by you.

• Experience with production PostgreSQL and Redis under real traffic, focusing on slow queries, connection limits, and zero-downtime schema migrations.

• Familiarity with production message brokers; NATS or similar.

• Experience implementing Cloudflare or similar edge solutions for public services.

• Capable of troubleshooting live incidents across application logs, metrics, databases, and edge services.

• AI-forward mindset with fluency in AI technologies.

• Ability to operate independently within existing systems.

• Strong written communication skills for both technical and non-technical audiences.

• Collaborative working style.

• Nice to have: AI-consumable documentation, self-service infrastructure pathways, security expertise, cloud cost optimization, experience with live services or gaming, and solo or near-solo infrastructure management experience.


🏝️ Benefits

• Unlimited vacation.

• Annual team offsite.

• Complete work equipment provided.

• Budget allocated for experimentation with any AI tools.

• Access to all company metrics and internal discussions.

• Asynchronous culture that promotes deep work without performative meetings.

People also viewed

Yopeso16 hours ago

Reliability Engineer / DevOps – Database Platform

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Orion Innovation16 hours ago

DevOps

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Climavision16 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $170k/year
ApplyView job
Parasail16 hours ago

Senior Site Reliability Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NICE16 hours ago

Cloud Operations Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Humana17 hours ago

Senior DevOps Engineer

US flagDistrict of Columbia, +2 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$106.9k – $147k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers