
Staff Engineer, Core β MLOps
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Portugal.
β’ Design and enhance the control plane, encompassing service registry, schema registry, SLO enforcement, and CLI tools.
β’ Create and establish the context plane to consolidate operational signals and develop automated feedback loops for self-healing maintenance.
β’ Take ownership of the multi-language service chassis and golden path, including Java and Python client libraries, workload specifications, Helm charts, and deployment pipelines.
β’ Define inter-service contracts using gRPC and Protocol Buffers, along with API gateway transcoding, versioning policies, and schema evolution guidelines.
β’ Manage the platform substrate across Kubernetes, Terraform, ingress infrastructure, Kafka, billing pipelines, Valkey, and database modernization projects.
β’ Lead Requests for Discussion regarding workflow orchestration, gateway orchestration, multi-cluster routing, and automated failover strategies.
β’ Establish SLOs and error budgets, implement health-aware traffic isolation, and facilitate automated weighted canary deployments.
β’ Engage in shared infrastructure on-call rotation and spearhead incident post-mortems.
β’ Mentor engineers across various teams, review system proposals, and promote engineering best practices.
β’ Act as the technical lead for the Core & MLOps squad, influencing architecture and release velocity across five adjacent squads.
β’ Over 10 years of experience in developing scalable distributed backend systems.
β’ Proven history of creating internal platforms or core libraries that are widely utilized by engineering teams.
β’ Expertise in Java, particularly with reactive frameworks like Vert.x or Netty.
β’ Strong proficiency in Python.
β’ Extensive experience with gRPC and Protocol Buffers.
β’ Background in managing schema evolution and ensuring backward compatibility in mission-critical settings.
β’ Practical experience with Kubernetes in production environments at scale.
β’ Familiarity with Terraform and event streaming technologies such as Kafka.
β’ Experience designing automated telemetry pipelines, materialized views, or feature stores that adapt system behavior based on real-time data.
β’ Demonstrated success in defining SLOs/SLIs, analyzing blast radius, and building fault-tolerant systems grounded in stringent service contracts.
β’ Excellent technical writing abilities.
β’ Strong interpersonal and written communication skills suited for a globally distributed, remote-first workplace.
β’ Experience with Java 21, Helm, CircleCI, OCI, GCP, Hetzner, Servers.com, Confluent Kafka, BigQuery, Valkey, MySQL/PostgreSQL, HBase, Google Pub/Sub, Prometheus, Grafana, Loki, and OpenTelemetry.
β’ Bonus: familiarity with Temporal, DBOS, model serving, performance monitoring, drift detection, SPIRE, mTLS, Cilium, Istio, Envoy, CLIs, SDKs, project generators, web scraping/crawling, or contributions to open-source projects.
β’ Freedom and flexibility to work from your optimal location.
β’ Flexible working hours.
β’ Opportunities to attend conferences and connect with team members worldwide.
β’ Engage with cutting-edge open source technologies and tools.
β’ A remote-first culture.
β’ Continuous innovation and challenging technical tasks.
β’ A global community fostering collaboration with distributed systems engineers and data specialists across the globe.
NationsBenefits
QAVION GROUP
Weflow | getweflow.com
Travoom
Get handpicked remote jobs straight to your inbox weekly.