
Senior DevOps Engineer
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in India.
• Transform manual runbooks into event-driven automation solutions.
• Design, create, and implement production-quality software utilizing Python and JavaScript/TypeScript.
• Assess blast radius and establish least-privilege IAM policies.
• Ensure that automated actions are both traceable and auditable.
• Provide support for core database and streaming platforms in a production environment.
• Engage in the team's production on-call rotation.
• Develop event-driven automations leveraging Python/TypeScript and AWS.
• Convert significant operational runbooks into secure, self-service or auto-triggered systems.
• Create standardized deployment patterns.
• Spearhead the adoption of LLM-enhanced development tools with code-review safeguards.
• Take ownership of the automation roadmap for the Data Infrastructure team.
• Collaborate with SRE, Security, and product teams to promote event-driven, self-healing systems.
• Bachelor’s degree in a relevant field, or equivalent professional experience.
• Over 5 years of professional software engineering experience, including significant time spent building or managing production systems at a SaaS or cloud service provider.
• Experience in 24x7 production operations for a highly available SaaS or cloud environment, including on-call duties.
• Proficient programming skills in Python and JavaScript/TypeScript; capable of managing services from design through to production.
• Extensive hands-on experience with AWS infrastructure, particularly Lambda, EventBridge, IAM, CloudWatch, and CloudTrail.
• Proficient in deploying, managing, and troubleshooting production databases and distributed data stores such as PostgreSQL, MySQL, Redis, Elasticsearch, DynamoDB, or Kafka.
• Experience in designing event-driven architectures and automation triggered by monitoring.
• Solid practical knowledge of Kubernetes in a production setting, encompassing deployments, RBAC, controllers, and safe operational practices.
• Experience with Prometheus and alert-driven workflows.
• Strong understanding of security principles, including least-privilege IAM design, secrets management, and blast-radius analysis.
• Proven capability in building for auditability, including structured logs, correlation IDs, and immutable change records.
• Demonstrated experience utilizing LLM-enhanced development tools like Cursor and Claude, with responsible output-review methodologies.
• Familiarity with CI/CD processes such as GitHub Actions or Jenkins.
• Strong foundational knowledge of systems, networking, and troubleshooting distributed systems.
• Excellent written communication skills, with the ability to translate complex operational incidents into clear runbooks and code.
• Ability to work independently across time zones and collaborate effectively with a globally distributed team.
• Preferred: experience in reducing manual toil at scale, resulting in measurable decreases in alerts, tickets, or MTTR.
• Preferred: background in SRE, DevOps, or platform engineering within a large multi-tenant SaaS environment.
• Preferred: familiarity with Terraform and GitOps methodologies.
• Preferred: experience with Apache Airflow and Kafka.
• Comprehensive medical insurance for employees and their dependents.
• Accident insurance coverage.
• Term life insurance provided.
• All insurance premiums covered by SailPoint.
• Company-sponsored health check-ups for employees.
• Discounted rates for health check-ups for dependents.
• Annual performance-based bonus.
• Restricted stock units available at certain levels.
• 24 vacation days each year.
• 10 public holidays.
• Flexible work hours.
ICF
Point Wild (Formerly Pango Group)
Point Wild (Formerly Pango Group)
Sumsub
Get handpicked remote jobs straight to your inbox weekly.