
SRE Monitoring Platform Software Engineer, Entry Level
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in California, +1 more state.
• Contribute to Bitdeer's NeoCloud SRE platform for monitoring, securing, and managing a multi-region GPU rental fleet.
• Develop collection agents, as well as stores for metrics, logs, traces, and profiles, along with enrichment services and collection monitors.
• Write code for ingestion, querying, and storage paths.
• Assist in the development of alerting, correlation, and SLO frameworks; implement and optimize default alert rules.
• Contribute to topology services, cluster health rollups, and OSS-SRE-tool collection plugins for platforms such as Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay.
• Aid in the creation of remediation actuators, orchestration/workflow components, inspection probes, and job schedulers.
• Instrument services using metrics, logs, and traces with OpenTelemetry.
• Create dashboards and draft actionable on-call runbooks.
• Develop unit, integration, and contract tests for delivered components.
• Engage in chaos and soak tests led by senior engineers.
• Build components from design through production utilizing GitOps and the CI/CD release pipeline.
• Achieve declared SLOs and maintain systems free from drift.
• Operate the systems you build under guidance from senior engineers.
• Participate in on-call duties as a shadow before assuming primary responsibility.
• Aim to independently deliver components and manage a sub-context within 12 months.
• 0–2 years of experience in software engineering; recent graduates with strong projects or internships are encouraged to apply.
• Strong understanding of Go, Python, Java, or Rust; Go is preferred.
• Capability to write clean, tested, and readable code and articulate design decisions.
• Knowledge of data structures, algorithms, concurrency, TCP/HTTP networking, and operating system fundamentals.
• Understanding of distributed systems concepts, including idempotency, retries, back-pressure, caching, and eventual consistency.
• Practical experience with Prometheus, Grafana, Loki, or similar observability tools.
• Proficiency in writing basic PromQL queries and instrumenting services.
• Familiarity with Linux, shell commands, system logs, and common debugging tools.
• Basic knowledge of Kubernetes, including Pods, Services, and Deployments; experience running applications on Kubernetes is a plus.
• Experience with Git, branching, pull requests, and CI pipelines like GitHub Actions or GitLab CI.
• Discipline in unit and integration testing.
• Proficiency in written and spoken English.
• Curiosity and eagerness to learn about GPU/AI infrastructure, AIOps, distributed systems, and observability.
• Nice-to-have: internship or project experience in monitoring, observability, telemetry pipelines, or platform/SRE tooling.
• Nice-to-have: familiarity with GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray.
• Nice-to-have: exposure to AIOps/ML-related tools.
• Nice-to-have: contributions to open-source observability or cloud-native projects.
• Mentorship from senior and principal engineers.
• Opportunity to learn the complete observability stack at a production scale.
• Participation in on-call roles as a shadow prior to taking on primary responsibilities.
• Commitment to equal employment opportunities and non-discrimination.
Shield AI
Netflix
Travoom
HeroSoftware GmbH - Shopify Apps
Get handpicked remote jobs straight to your inbox weekly.