
DevOps Engineer, AI/ML Infrastructure
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in New York.
• Assist in the maintenance and support of a scalable, highly available, and secure cloud infrastructure.
• Provision and oversee cloud resources utilizing Infrastructure as Code.
• Implement best practices for cloud security, including IAM, role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.
• Provide support for containerized environments and orchestration platforms.
• Integrate DevSecOps principles throughout infrastructure and deployment workflows.
• Engage in disaster recovery planning, testing, and recovery initiatives.
• Maintain and optimize CI/CD pipelines that facilitate application and ML model deployments.
• Enhance deployment reliability while supporting zero-downtime deployment strategies.
• Automate configuration management, infrastructure provisioning, and routine operational tasks.
• Diagnose deployment and pipeline issues, implementing measures to prevent recurrence.
• Create scripts and automation to minimize manual tasks.
• Design, deploy, operate, and secure infrastructure that supports AI and agentic products.
• Manage and scale ML platform infrastructure, including Databricks clusters, job compute, ML pipelines, and Model Serving endpoints.
• Oversee production model-serving infrastructure, compute capacity, provisioned throughput, and autoscaling.
• Monitor model drift, data quality, inference performance, and serving health.
• Maintain monitoring, logging, metrics, and alerting systems using tools like Prometheus, Grafana, Coralogix, and CloudWatch.
• Assist in incident response and conduct Root Cause Analysis for infrastructure and deployment challenges.
• Collaborate with Software Engineers, ML Engineers, QA, Software Engineers in Test, IT Security, and Product Support.
• Respond to requests from engineering and Product Support teams.
• Keep accurate internal technical and operational documentation.
• Work collaboratively with international teams across various time zones.
• Over 3 years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-centric role.
• More than 2 years of hands-on experience with AWS services including EC2, S3, RDS, Lambda, IAM, VPC, SQS, and API Gateway, or comparable services.
• At least 2 years of experience with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.
• Proven experience in building and maintaining CI/CD pipelines using GitLab CI/CD or Jenkins.
• Practical experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar tools.
• Strong Linux system administration and troubleshooting capabilities.
• Solid understanding of networking principles, including routing, load balancing, and network security.
• Proficient in scripting languages such as Python, Bash, or similar.
• Practical experience with AI tools aimed at enhancing engineering workflows, automation, troubleshooting, or agentic applications.
• Experience in supporting data, ML, or other compute-intensive production workloads.
• Familiarity with Google Cloud is advantageous.
• Knowledge of Helm and service mesh technologies like Istio, Linkerd, or Traefik would be beneficial.
• Exposure to serverless and event-driven architectures is a plus.
• Familiarity with cloud and infrastructure security practices, vulnerability management, and tools such as Nessus, Prowler, Trivy, or firewalls is valuable.
• Understanding of security standards, compliance requirements, and cloud security best practices is a plus.
• Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, or Grafana is favorable.
• FinOps experience would be a plus.
• Experience with API gateways or API management platforms such as Kong or Apigee is beneficial.
• Familiarity with MLOps platforms and practices, particularly Databricks, model serving, ML pipelines, and model monitoring, would be advantageous.
• 15 days of vacation.
• Floating and company holidays.
• Wellness benefits.
• Paid parental leave.
• Comprehensive medical insurance.
• Dental insurance.
• Vision insurance.
• Life insurance.
• Disability insurance.
• 401(k) plan.
• Work from home stipend.
• Access to a learning platform, training, and tools.
• Flexible work model combining in-person and remote work.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.