
Senior DevOps Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in New York.
• Lead efforts to implement and uphold best practices for data streaming, processing, analytics, and monitoring infrastructure.
• Deploy and oversee services on Kubernetes-based platforms such as Amazon EKS and Google Kubernetes Engine (GKE).
• Provision and manage cloud infrastructure using Terraform, adhering to best practices in security, scalability, and cost efficiency.
• Maintain and enhance CI/CD pipelines utilizing Jenkins, ArgoCD, and GitHub Enterprise Actions to facilitate automated deployments and testing.
• Collaborate with cloud-native data services including AWS Kinesis, AWS Glue, Google Dataflow, Google Pub/Sub, BigQuery, and BigTable.
• Familiarity with workflow orchestration tools such as Apache Airflow and Google Cloud Composer.
• Develop and sustain automation scripts and tools using Python to support DevOps processes.
• Monitor system performance, resolve issues, and implement proactive solutions to improve reliability and efficiency.
• Adopt SRE practices to enhance service reliability, scalability, and cost-effectiveness.
• Analyze and optimize cloud expenses, identifying opportunities for improvement and implementing cost-saving measures.
• Ensure adherence to security policies and best practices within cloud environments.
• Promote the adoption of company standards and guide data teams to align with optimal DevOps and SRE practices.
• Collaborate with cross-functional teams to refine development workflows and infrastructure.
• Over 7 years of experience in a DevOps, Site Reliability Engineering, or Cloud Infrastructure position.
• Extensive experience with AWS and GCP data services, including Kinesis, Glue, Pub/Sub, and Dataflow.
• Proficiency in deploying and managing workloads on Kubernetes (EKS/GKE) within production settings.
• Practical experience with Infrastructure-as-Code (IaC) using Terraform.
• Expertise in managing CI/CD pipelines utilizing Jenkins, ArgoCD, and GitHub Enterprise Actions.
• Programming proficiency in Python for automation and scripting tasks.
• Experience with observability and monitoring tools (e.g., Prometheus, Grafana, Datadog, or CloudWatch).
• Strong grasp of SRE principles, including performance monitoring, incident response, and reliability engineering.
• Experience with cost optimization strategies for cloud infrastructure.
• Self-motivated and proactive, with a strong capability to influence and drive change across multiple teams.
• Ability to collaborate effectively in an agile environment and support various teams.
• Holistic mind, body, and lifestyle programs designed for overall well-being.
• Comprehensive benefits.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.