SRE Monitoring Platform Software Engineer – Early Career, Temporary

Posted Aug 27

This is a fully remote position, open to applicants in California, +1 more state.

πŸ“‹ Description

β€’ Contribute to the NeoCloud SRE platform, which is a multi-region system managing a GPU rental fleet across various data centers.

β€’ Construct well-defined components from the design phase through to production code with guidance from senior engineers.

β€’ Deploy code via GitOps and CI/CD release pipelines while adhering to Plugin Framework standards and defined SLOs.

β€’ Develop components for collection and storage, including ingestion, query, storage-path, enrichment, and collection-monitor code.

β€’ Assist in the creation of alerting, correlation, and SLO frameworks; implement and refine default alert rules.

β€’ Create topology and cluster-health services along with collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay.

β€’ Aid in the development of remediation actuators, orchestration/workflow components, inspection probes, and job schedulers.

β€’ Instrument services with metrics, logs, and traces utilizing OpenTelemetry.

β€’ Design dashboards and draft actionable on-call runbooks.

β€’ Write unit, integration, and contract tests for the components delivered.

β€’ Engage in chaos and soak tests led by senior engineers.

β€’ Operate the developed systems under supervision and participate in on-call duties as a shadow before assuming primary responsibilities.

β€’ Aim to independently deliver components and manage a sub-context within a 12-month timeframe.


⛳️ Requirements

β€’ 0–2 years of experience in software engineering; recent graduates with strong projects or internships are encouraged to apply.

β€’ Strong foundational knowledge in at least one programming language: Go (preferred), Python, Java, or Rust.

β€’ Capability to write clean, tested, and readable code while articulating design decisions.

β€’ Understanding of data structures, algorithms, concurrency, TCP/HTTP networking, and operating system principles.

β€’ Familiarity with distributed systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency.

β€’ Practical experience with monitoring and observability tools such as Prometheus, Grafana, or Loki; proficiency in writing basic PromQL and instrumenting a service.

β€’ Knowledge of Linux, shell scripting, system logs, and common debugging tools.

β€’ Fundamentals of Kubernetes, including Pods, Services, and Deployments; experience running applications on Kubernetes.

β€’ Understanding of Git and CI fundamentals, including branching, pull requests, and CI pipeline usage.

β€’ Regular practice of writing unit and integration tests.

β€’ Strong verbal and written communication skills in English.

β€’ Curiosity and a desire to learn about GPU/AI infrastructure, AIOps, distributed systems, and observability.

β€’ Nice-to-have: internship or project experience in monitoring/observability, telemetry pipelines, or platform/SRE tooling.

β€’ Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray.

β€’ Nice-to-have: familiarity with AIOps or ML-adjacent tools.

β€’ Nice-to-have: contributions to open-source observability or cloud-native projects.


🏝️ Benefits

β€’ Mentorship from senior and principal engineers.

β€’ Guided involvement in on-call, starting as a shadow before taking on primary responsibilities.

β€’ Exposure to production-scale observability, GPU/AI infrastructure, AIOps, distributed systems, and cloud-native technologies.

β€’ Opportunity to contribute to greenfield systems utilizing an established Plugin Framework, GitOps pipeline, and SLO framework.

β€’ Equal employment opportunities.

People also viewed

RTX1 day ago

Principal Software Engineer – Air Traffic Solutions

US flagMaine OnlyFull-timeFull-stack Engineer$107.5k – $204.5k/year
ApplyView job
NVIDIA1 day ago

Senior System Software Engineer, Software-Defined Networking

US flagCalifornia, +4 more statesFull-timeFull-stack Engineer$224k – $356.5k/year
ApplyView job
Infios1 day ago

Senior Software Engineer

MX flagMexico, +1 more countryFull-timeFull-stack Engineer
ApplyView job
Gramian Consulting2 days ago

Software Engineer – Licensing, AI Training

EG flagEgypt, +3 more countriesFull-timeFull-stack Engineer$100/year
ApplyView job
Gramian Consulting2 days ago

Software Engineer – Licensing, AI Training

IN flagIndia, +5 more countriesFreelanceFull-stack Engineer$100/year
ApplyView job
Gramian Consulting2 days ago

Software Engineer – License your Git Repositories for AI Training

CA flagCanada, +5 more countriesFull-timeFull-stack Engineer$100/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers