
Engineer III – TechOps CICD SRE, Reliability Focused
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in India.
• Develop software and systems to oversee platform infrastructure and applications.
• Assist with Crowdstrike’s main CI/CD build tools.
• Create automation solutions for service deployment.
• Monitor availability and maintain a comprehensive view of system health.
• Enhance the reliability, quality, and serviceability of systems.
• Engage in architectural design of highly available services at an enterprise level.
• Collaborate with internal customers to identify needs and create solutions.
• Lead the Incident Response and Production Readiness Review (PRR) initiatives within our organization.
• Collect and evaluate metrics from both operating systems and applications to aid in performance tuning and root cause analysis.
• Conduct resource, capacity, and license forecasting.
• Collaborate with and promote a mentorship mindset among fellow Engineers.
• Design, implement, and sustain full-stack observability (Datadog, Grafana stack, LogScale, Jaeger, New Relic, OpenTelemetry, Prometheus/Thanos, etc.) to fulfill production reliability requirements.
• Optimize for high-cardinality workloads and standardize instrumentation.
• Propel the shift towards AI-native observability to enhance system resilience and achieve zero-toil operations.
• Automate alert correlation, anomaly detection, and root cause analysis via AIOps integration.
• Train internal AI models using runbooks and incident data to enable self-healing frameworks.
• Maintain code quality and secure development standards (SonarQube, OWASP) throughout the engineering lifecycle.
• Regularly review SLI/SLO definitions and error budgets to ensure they accurately reflect real user impact.
• Over 5 years of experience in a large-scale production environment.
• On-Premise & Cloud proficiency in deploying, scaling, and managing CI/CD tools such as Bazel, Github Actions, and Jenkins.
• Familiarity with IaC Provisioning tools including Ansible, Chef, Puppet, Salt, and Terraform.
• Experienced in Source Code Management services like Bitbucket, Gitlab, and Github.
• Knowledge of Monitoring and Observability tools including Open-source Prometheus/Grafana, Datadog, Honeycomb, and New Relic.
• Proven experience in deploying applications on Kubernetes at scale.
• Demonstrated ability to collaborate effectively with both local and remote teams.
• Must possess strong attention to detail and the capacity to make timely, informed decisions.
• Ability to balance short-term and long-term objectives.
• Show initiative and self-learning capabilities in a fast-paced and rapidly changing environment.
• Security-first mindset with a general understanding of cybersecurity principles.
• Experience integrating AI into existing workflows.
• Strong programming skills in one or more languages such as Python, Go, TypeScript, Java, or C#.
• Industry-leading compensation and equity awards.
• Comprehensive wellness programs for physical and mental health.
• Competitive vacation and holiday policies for rest and recharge.
• Paid parental and adoption leave.
• Professional development opportunities available to all employees, regardless of role or level.
• Employee Networks, local neighborhood groups, and volunteer opportunities to foster connections.
• A vibrant office culture featuring world-class amenities.
Cribl
Flock Safety
Pear Tree.
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.