
SRE Platform Software Engineer β Early Career, Temporary
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
β’ Develop, test, and implement features for essential SRE platform components, which encompass collection agents, telemetry pipelines, alert systems, and cluster health services.
β’ Execute infrastructure and application updates utilizing GitOps methodologies, declarative configurations, and automated CI/CD pipelines.
β’ Monitor, evaluate, and enhance system metrics, logs, and traces to maintain high availability across GPU infrastructure.
β’ Engage in learning and participate in the on-call rotation for services managed by the team.
β’ Create comprehensive runbooks and incident post-mortem reports.
β’ Produce unit, integration, and end-to-end tests to ensure platform reliability prior to production deployment.
β’ Collaborate with senior engineers to transform architect-approved designs into production-ready code.
β’ Assist in operating critical services utilized by other teams, cloud groups, and tenants.
β’ Ensure infrastructure remains drift-free through the Plugin Framework and established operational practices.
β’ Bachelorβs degree in Computer Science, Computer Engineering, or a related technical field (or equivalent practical experience/internships).
β’ 0β2 years of practical software development experience.
β’ Proficient in Go (preferred), Java, or Rust.
β’ Strong scripting capabilities in Python or Bash.
β’ Solid foundational understanding of data structures, algorithms, object-oriented design, and concepts of distributed systems, including APIs, concurrency, and networking basics.
β’ Hands-on experience with Docker, Kubernetes, and Linux fundamentals through coursework, personal projects, open-source contributions, or internships.
β’ Experience in writing unit and integration tests for personal code.
β’ Strong technical writing abilities for documenting design decisions, runbooks, and clear Pull Request descriptions.
β’ Previous internship or project experience with Kubernetes Operators, Helm, or GitOps tools such as ArgoCD or Flux is a significant advantage.
β’ Familiarity with time-series databases or observability tools such as Prometheus, OpenTelemetry, Grafana, or Loki is a strong plus.
β’ Basic knowledge of hardware, GPU/AI infrastructure like NVIDIA DCGM and CUDA, or concepts related to high-performance computing is a strong plus.
β’ Familiarity with infrastructure-as-code tools such as Terraform or Ansible is a plus.
β’ Full-time employment.
β’ Mentored on-call rotation and early-career development in collaboration with senior engineers.
β’ Equal employment opportunities and a commitment to non-discrimination.
Sigma Software Group
Collectly
Allata
Get handpicked remote jobs straight to your inbox weekly.