
Software Engineer – Infrastructure
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California.
• Take ownership of the platform, which includes GCP, Kubernetes, Temporal, the GPU fleet, along with deployment and rollback mechanisms.
• Engage in the on-call rotation while enhancing its efficiency and responsiveness.
• Oversee GPU capacity, training and inference pipelines, and ensure the reliability and cost-effectiveness of production model-serving.
• Make informed decisions regarding infrastructure costs and manage metering and attribution as inference usage increases.
• Administer identity and access management, secrets management, enforce least-privilege boundaries, and ensure supply-chain integrity.
• Develop clear infrastructure-as-code, runbooks, in-repo documentation, and observability tools.
• Enhance tooling, standards, testing protocols, observability, and release practices.
• Offer architectural guidance, conduct thoughtful reviews, provide mentoring, and maintain clear documentation.
• Collaborate with engineering teams to understand their requirements and develop the platform roadmap.
• A minimum of 8 years of experience in building and operating production distributed systems, or equivalent server-side engineering with a strong focus on infrastructure.
• Ability to effectively utilize agents and critically assess the application of AI.
• Experience working in environments where failures are costly.
• Proven track record of carrying a pager, managing incidents, and executing system rollbacks.
• Proficient in using SLOs and error budgets as operational tools.
• Hands-on experience with a major cloud provider and Kubernetes in a production environment.
• Familiarity with infrastructure-as-code as a standard practice.
• Demonstrated ownership of an architecture or migration process from planning through to launch.
• Capability to identify unassigned tasks, define their scope, secure buy-in, and deliver results.
• Proficient in building minimal reproductions, analyzing logs, and crafting targeted checks.
• Valuable experience in GPU and ML infrastructure.
• Relevant experience in production security engineering, IAM, secrets management, supply chain, and least privilege practices.
• Insightful experience in cloud cost modeling, commitment strategies, reservations, and unit economics.
• Beneficial experience with CI/CD at a monorepo scale and in developer environments.
• Advantageous experience with media, video, or GPU-supported workloads.
• Relevant experience working within small teams while managing significant responsibilities.
• Equity.
• Comprehensive healthcare package.
• 401k matching program.
• Catered lunches.
• Flexible vacation policy.
• Year-round opportunities for in-person collaboration for remote employees.
Rocket Money (formerly Truebill)
Leidos
Precise Software Solutions, Inc.
Coinbase
Get handpicked remote jobs straight to your inbox weekly.