
SRE Monitoring Platform Software Engineer β Early Career, Temporary
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in California, +1 more state.
β’ Contribute to the NeoCloud SRE platform, which is a multi-region system managing a GPU rental fleet across various data centers.
β’ Construct well-defined components from the design phase through to production code with guidance from senior engineers.
β’ Deploy code via GitOps and CI/CD release pipelines while adhering to Plugin Framework standards and defined SLOs.
β’ Develop components for collection and storage, including ingestion, query, storage-path, enrichment, and collection-monitor code.
β’ Assist in the creation of alerting, correlation, and SLO frameworks; implement and refine default alert rules.
β’ Create topology and cluster-health services along with collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay.
β’ Aid in the development of remediation actuators, orchestration/workflow components, inspection probes, and job schedulers.
β’ Instrument services with metrics, logs, and traces utilizing OpenTelemetry.
β’ Design dashboards and draft actionable on-call runbooks.
β’ Write unit, integration, and contract tests for the components delivered.
β’ Engage in chaos and soak tests led by senior engineers.
β’ Operate the developed systems under supervision and participate in on-call duties as a shadow before assuming primary responsibilities.
β’ Aim to independently deliver components and manage a sub-context within a 12-month timeframe.
β’ 0β2 years of experience in software engineering; recent graduates with strong projects or internships are encouraged to apply.
β’ Strong foundational knowledge in at least one programming language: Go (preferred), Python, Java, or Rust.
β’ Capability to write clean, tested, and readable code while articulating design decisions.
β’ Understanding of data structures, algorithms, concurrency, TCP/HTTP networking, and operating system principles.
β’ Familiarity with distributed systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency.
β’ Practical experience with monitoring and observability tools such as Prometheus, Grafana, or Loki; proficiency in writing basic PromQL and instrumenting a service.
β’ Knowledge of Linux, shell scripting, system logs, and common debugging tools.
β’ Fundamentals of Kubernetes, including Pods, Services, and Deployments; experience running applications on Kubernetes.
β’ Understanding of Git and CI fundamentals, including branching, pull requests, and CI pipeline usage.
β’ Regular practice of writing unit and integration tests.
β’ Strong verbal and written communication skills in English.
β’ Curiosity and a desire to learn about GPU/AI infrastructure, AIOps, distributed systems, and observability.
β’ Nice-to-have: internship or project experience in monitoring/observability, telemetry pipelines, or platform/SRE tooling.
β’ Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray.
β’ Nice-to-have: familiarity with AIOps or ML-adjacent tools.
β’ Nice-to-have: contributions to open-source observability or cloud-native projects.
β’ Mentorship from senior and principal engineers.
β’ Guided involvement in on-call, starting as a shadow before taking on primary responsibilities.
β’ Exposure to production-scale observability, GPU/AI infrastructure, AIOps, distributed systems, and cloud-native technologies.
β’ Opportunity to contribute to greenfield systems utilizing an established Plugin Framework, GitOps pipeline, and SLO framework.
β’ Equal employment opportunities.
RTX
NVIDIA
Infios
Gramian Consulting
Get handpicked remote jobs straight to your inbox weekly.