
Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States, +1 more country.
• Design, implement, and oversee tools utilized for the development, deployment, monitoring, and support of cloud infrastructure and services.
• Establish and monitor SLIs and SLOs, employing error budgets to balance reliability initiatives with feature velocity.
• Monitor and troubleshoot cloud resources and applications through the observability stack.
• Engage in an on-call rotation and collaborate with development and IT teams to resolve incidents.
• Automate and enhance processes to increase the efficiency and reliability of cloud operations.
• Monitor and analyze cloud resource consumption to uncover cost-saving opportunities.
• Construct and sustain infrastructure-as-code for the Azure environment utilizing Bicep/OpenTofu.
• Facilitate the transition from manual portal-driven modifications to GitOps-based workflows.
• Assist in the deployment and configuration of InRule products and solutions.
• Propose changes in application architecture or design to enhance security, performance, and efficiency.
• Implement Azure security measures, including access management, secure baselines, and monitoring.
• Support least-privilege access, just-in-time elevation with Azure PIM, and initiatives for detecting infrastructure drift.
• Aid in the onboarding or tuning of SAST, DAST, and SCA security scanning in CI/CD pipelines.
• Design and deploy workflows for tracking release changes, vulnerability triage, and incident response.
• A minimum of 5 years of experience in cloud operations, site reliability engineering, or DevOps.
• Robust understanding of core cloud services, including compute, serverless, managed databases, storage, and secrets management on AWS, Azure, or GCP.
• Proficiency in PowerShell and other scripting or programming languages such as SQL, Python, C#, or JavaScript.
• Experience with infrastructure-as-code tools like OpenTofu, Bicep/ARM, or Pulumi.
• Direct experience with at least one major observability platform: Datadog, New Relic, ELK, Prometheus/Loki/Grafana.
• Familiarity with OpenTelemetry-based instrumentation.
• Strong DevOps principles and CI/CD experience with GitHub Actions, Azure DevOps, or similar tools.
• Experience with Linux operating systems and tools such as Docker and Kubernetes.
• Willingness to take responsibility for security-related tasks.
• Experience with Azure Cloud, including Functions, Container Solutions, SQL Database, Storage, and Key Vault.
• Practical experience with Azure security tools, including Entra ID, Defender for Cloud, Sentinel, and Azure Policy.
• Background in operating within a regulated environment such as SOC 2, ISO 27001, or HIPAA.
• Knowledge of networking, DNS, VPC, and database operations and concepts.
• Exposure to Windows Server and IIS configuration and maintenance.
• Experience with configuration management tools such as Ansible, Puppet, or Chef.
• Must be authorized to work in the United States without current or future sponsorship, or be located in Sweden for the Sweden position.
• Competitive compensation and benefits package.
• Flexible working environment.
• Opportunities for professional growth within a scaling SaaS organization.
• Collaborative culture with strong partnerships across Support, Engineering, Product, and Customer Success.
• Chance to build and shape a premium support function that delivers measurable customer impact.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
Thumbtack
Get handpicked remote jobs straight to your inbox weekly.