
Site Reliability Engineer – Public Sector
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in United States.
• Deploy, operate, and maintain Blitzy's self-hosted platform within a secure cloud environment controlled by the customer.
• Take ownership of Kubernetes-based deployments, including releases, upgrades, capacity planning, and performance benchmarking for compute-intensive AI tasks.
• Design and sustain observability that encompasses logging, metrics, tracing, and alerting within the security perimeter set by the customer.
• Act as Blitzy's technical representative on account, collaborating with customer infrastructure, security, and governance teams.
• Assist with provisioning, reviews, documentation, and operational escalations.
• Manage sensitive customer data in accordance with security requirements and advocate for best practices in security.
• Incorporate lessons learned into Blitzy's product and infrastructure development roadmap.
• Facilitate customer onboarding, deployment strategy, monitoring, alerting, incident response, SLOs, and capacity validation.
• U.S. citizenship is required, along with the ability to complete a customer background and badging process.
• A minimum of 3 years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
• Strong expertise in Kubernetes and container orchestration.
• Practical experience deploying software in customer-controlled or restricted environments.
• Experience operating in isolated, restricted, or highly regulated network settings.
• Hands-on experience with infrastructure-as-code tools such as Terraform, Pulumi, or their equivalents.
• Familiarity with at least one major cloud platform.
• In-depth knowledge of observability tools, incident management, and on-call protocols.
• Strong scripting and automation capabilities in Python, Go, Bash, or similar languages.
• Excellent communication skills for effective collaboration with customer engineering, security, and governance stakeholders.
• Experience deploying or managing software in government-accredited or similarly certified cloud environments is advantageous.
• Understanding of sensitive data management and security frameworks in regulated industries is a plus.
• Experience in supporting AI/ML workloads or related infrastructure is beneficial.
• Previous experience in forward-deployed, residency, or embedded-engineer roles is a plus.
• Experience in a high-growth startup environment is a plus.
• Bonus aligned with experience.
• Equity based on experience.
• Remote work options.
• Occasional travel for essential customer workshops.
• Opportunity to collaborate closely with top-tier engineers.
• Direct impact on architectural decisions.
• Equal opportunity employer dedicated to fostering a diverse and inclusive team.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.