
Platform Engineer
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in Spain, +4 more countries.
• Take ownership of a designated set of platform services, ensuring their reliability within established service levels.
• Manage the observability platform, facilitate team onboarding, oversee cost and capacity, and uphold alerting standards.
• Administer GitLab and the CI runner fleet, covering upgrades, capacity management, access control, backups, and restore exercises.
• Support additional services through production monitoring and the creation of runbooks.
• Implement requested services from the ground up using suitable designs, infrastructure as code, monitoring solutions, backups, and documentation.
• Address developer inquiries regarding access, onboarding, pipeline issues, exporters, and dashboards; transform recurring requests into self-service solutions.
• React to incidents by diagnosing issues, minimizing impacts, safely restoring services, conducting root-cause analyses and post-mortems, and enacting prevention or detection enhancements.
• Deliver changes as code, ensuring they undergo review in merge requests, while planning and validating each modification.
• Create runbooks, onboarding resources, maintenance updates, and actionable status reports for engineers outside the team.
• Collaborate with AI agents by assigning tasks for data collection and drafting, reviewing their outputs, and documenting insights for the team.
• Partner with product and engineering teams to discuss scope, priorities, timelines, and platform requirements.
• Extensive experience at a senior level in infrastructure, platform, or site reliability engineering, with a track record of maintaining at least one production service.
• Proficiency in Linux systems administration and troubleshooting on both bare metal and virtual machines.
• Hands-on experience with Kubernetes in production, implemented through GitOps, including personal execution of cluster upgrades.
• Expertise in infrastructure as code using tools like Ansible and Terraform or OpenTofu, with changes reviewed in merge requests.
• Experience with GitLab administration and GitLab CI in a production environment, whether self-hosted or SaaS; similar proficiency with another CI system is acceptable.
• Familiarity with Prometheus and Grafana, including team operation, writing alert rules and dashboards, and understanding PromQL.
• Capability to craft technical documentation for engineers outside the team, including runbooks, notices, and responses to requests.
• Excellent communication and interpersonal abilities.
• Advanced experience using AI engineering assistants like Claude and Codex.
• Skill in providing context, segmenting tasks, designing agent workflows, and delegating plans for autonomous execution within set boundaries and permissions.
• Ability to articulate, debug, and test the resulting automation, ensuring generated commands, scripts, and conclusions are verified before production implementation.
• Proficiency in English at an upper-intermediate level or higher.
• Nice to have: experience with alerting design, SLOs, burn-rate alerts, and data-sized thresholds.
• Nice to have: familiarity with Kata Containers, Firecracker, or gVisor.
• Nice to have: experience with S3-compatible object storage operations like Ceph RGW.
• Nice to have: practical knowledge of AWS with real cost management.
• Nice to have: experience with self-hosted Sentry or other applications backed by Kafka, ClickHouse, and Redis under load.
• Nice to have: proficiency in Python or Go for developing exporters and small internal services.
• Emphasis on professional growth and development.
• Engaging and challenging projects.
• Fully remote work with flexible hours, enabling work from any location worldwide.
• Paid vacation of 24 days per year.
• 10 national holidays.
• Unlimited sick leave policy.
• Compensation for private medical insurance.
• Reimbursement for co-working and gym/sports expenses.
• Budget allocated for educational purposes.
• Opportunity to earn a reward for the most innovative idea that the company can patent.
EasyLlama - HR & Compliance Training For Modern Teams
Coinbase
Stitch Fix
Ambry Genetics
Get handpicked remote jobs straight to your inbox weekly.