Distinguished Engineer, Core DevOps

Posted 1 day ago

This is a fully remote position, open to applicants in United States, +2 more countries.

📋 Description

• Establish and continuously enhance the future technical direction for CI, CD, Plan, and source code experiences.

• Collaborate closely with the VP of Engineering on strategic planning, technical investments, and engineering priorities.

• Convert company-wide objectives for productivity, quality, and AI-native engineering into technical direction and iterative roadmaps for Core DevOps.

• Set measurable quality standards for availability, p95 and p99 latency, pipeline success rate, test reliability, data accuracy, and defect escape rate.

• Recognize systemic technical risks and develop prioritized plans to mitigate them.

• Lead intricate design decisions through coding, prototyping, and proofs of concept.

• Promote alignment on shared patterns, libraries, and established workflows.

• Create incremental migration, dual-run, rollback, and deprecation strategies for long-term systems.

• Integrate AI platform capabilities into Core DevOps and communicate requirements back to the AI organization.

• Act as an escalation point for complex or disputed technical decisions.

• Evaluate critical-path designs and merge requests, using reviews as teaching opportunities.

• Ensure designs are compatible across GitLab.com, GitLab Dedicated, and Self-Managed deployments while observing multi-tenant, compliance, and data-governance boundaries.

• Collaborate with Infrastructure, Security, and SRE on observability, debuggability, and graceful failure patterns.

• Communicate architectural constraints and opportunities to Product through roadmap language.

• Compose design documents, architecture narratives, and decision records.

• Mentor Principal and Staff Engineers and help elevate the senior technical hiring standards.

• Engage in the Incident Management on-call rotation and complete Interview Training for technical hiring processes.


⛳️ Requirements

• Over 10 years of experience in software engineering, including at least 4 years in a Staff, Principal, or equivalent senior technical leadership role.

• Profound expertise in AI and ML systems, encompassing large language models, agentic frameworks, and autonomous workflow design at a production scale.

• Demonstrated success in leading hands-on technical experimentation, including defining evaluation frameworks, conducting benchmarks, and translating insights into scalable architectural decisions.

• Strong foundation in scalable, multi-tenant distributed systems, covering service decomposition, fault tolerance, observability, and operational resilience.

• Experience in designing and implementing human-in-the-loop controls, safety guardrails, and responsible AI practices for production environments.

• Proven ability to mentor senior engineers and influence technical direction across multiple teams or divisions without direct authority.

• Strong capability to work efficiently in a fully remote, globally distributed organization with excellent written and asynchronous communication skills.


🏝️ Benefits

• Benefits to support your health, finances, and well-being

• Flexible Paid Time Off

• Team Member Resource Groups

• Equity Compensation & Employee Stock Purchase Plan

• Growth and Development Fund

• Parental Leave

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers