
Staff Site Reliability Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in California.
• Design, construct, and manage infrastructure that supports Circle’s blockchain platform at scale.
• Take responsibility for the reliability, performance, and automation of distributed blockchain systems across various public cloud environments.
• Operate and expand production blockchain infrastructure, including full nodes across Arc, Ethereum, Solana, Base, and other protocols.
• Create, build, and sustain highly available, secure Kubernetes platforms.
• Develop and enhance CI/CD pipelines, deployment workflows, and Infrastructure as Code.
• Innovate AI-driven infrastructure automation for Kubernetes lifecycle management, IaC analysis, blockchain cost optimization, and MCP integrations.
• Monitor, diagnose, and enhance distributed systems.
• Engage in follow-the-sun on-call rotation, incident response, and post-incident evaluations.
• Collaborate with protocol, product, and security teams on network launches, upgrades, and business requirements.
• Mentor engineers, disseminate SRE best practices, and contribute to architecture and reliability enhancements.
• Over 6 years of experience in SRE, DevOps, or Infrastructure Engineering.
• Experience in cloud-native or blockchain settings.
• Demonstrated technical leadership in the architecture and design of complex distributed systems.
• Experience in operating and troubleshooting blockchain nodes and distributed systems in a production environment.
• Extensive knowledge of blockchain protocols and their infrastructure requirements.
• In-depth understanding of Kubernetes internals and tuning for large-scale workloads.
• Strong experience with Kubernetes, including Helm, RBAC, operators, controllers, and observability tools.
• Proven experience in building CI/CD pipelines, containerized workloads, and implementing blue-green or canary releases.
• Proficiency in Infrastructure as Code using Terraform or Pulumi.
• Understanding of cloud networking fundamentals, including VPCs, DNS, load balancers, and secure connectivity.
• Proficient in Go or Python.
• Experience managing SQL databases and stateful production services.
• Familiarity with agentic automation strategies and AI tools in engineering processes.
• Knowledge of observability and chaos engineering methodologies.
• Experience mentoring engineers and implementing SRE best practices across teams.
• Strong engineering judgment, commitment to code quality, automated testing, and operational excellence.
• Ability to lead cross-functional technical projects and communicate effectively across engineering teams.
• Flexible work environment.
• Inclusive workplace culture.
• Opportunities for growth, continual challenges, and learning.
• Interview accommodations and support for candidates with disabilities.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.