DevOps Engineer, AI/ML Infrastructure

atCreatorIQRemoteUS flagNew YorkFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$106k – $127k/year

Posted 1 day ago

This is a fully remote position, open to applicants in New York.

📋 Description

• Assist in the maintenance and support of a scalable, highly available, and secure cloud infrastructure.

• Provision and oversee cloud resources utilizing Infrastructure as Code.

• Implement best practices for cloud security, including IAM, role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.

• Provide support for containerized environments and orchestration platforms.

• Integrate DevSecOps principles throughout infrastructure and deployment workflows.

• Engage in disaster recovery planning, testing, and recovery initiatives.

• Maintain and optimize CI/CD pipelines that facilitate application and ML model deployments.

• Enhance deployment reliability while supporting zero-downtime deployment strategies.

• Automate configuration management, infrastructure provisioning, and routine operational tasks.

• Diagnose deployment and pipeline issues, implementing measures to prevent recurrence.

• Create scripts and automation to minimize manual tasks.

• Design, deploy, operate, and secure infrastructure that supports AI and agentic products.

• Manage and scale ML platform infrastructure, including Databricks clusters, job compute, ML pipelines, and Model Serving endpoints.

• Oversee production model-serving infrastructure, compute capacity, provisioned throughput, and autoscaling.

• Monitor model drift, data quality, inference performance, and serving health.

• Maintain monitoring, logging, metrics, and alerting systems using tools like Prometheus, Grafana, Coralogix, and CloudWatch.

• Assist in incident response and conduct Root Cause Analysis for infrastructure and deployment challenges.

• Collaborate with Software Engineers, ML Engineers, QA, Software Engineers in Test, IT Security, and Product Support.

• Respond to requests from engineering and Product Support teams.

• Keep accurate internal technical and operational documentation.

• Work collaboratively with international teams across various time zones.


⛳️ Requirements

• Over 3 years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-centric role.

• More than 2 years of hands-on experience with AWS services including EC2, S3, RDS, Lambda, IAM, VPC, SQS, and API Gateway, or comparable services.

• At least 2 years of experience with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.

• Proven experience in building and maintaining CI/CD pipelines using GitLab CI/CD or Jenkins.

• Practical experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar tools.

• Strong Linux system administration and troubleshooting capabilities.

• Solid understanding of networking principles, including routing, load balancing, and network security.

• Proficient in scripting languages such as Python, Bash, or similar.

• Practical experience with AI tools aimed at enhancing engineering workflows, automation, troubleshooting, or agentic applications.

• Experience in supporting data, ML, or other compute-intensive production workloads.

• Familiarity with Google Cloud is advantageous.

• Knowledge of Helm and service mesh technologies like Istio, Linkerd, or Traefik would be beneficial.

• Exposure to serverless and event-driven architectures is a plus.

• Familiarity with cloud and infrastructure security practices, vulnerability management, and tools such as Nessus, Prowler, Trivy, or firewalls is valuable.

• Understanding of security standards, compliance requirements, and cloud security best practices is a plus.

• Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, or Grafana is favorable.

• FinOps experience would be a plus.

• Experience with API gateways or API management platforms such as Kong or Apigee is beneficial.

• Familiarity with MLOps platforms and practices, particularly Databricks, model serving, ML pipelines, and model monitoring, would be advantageous.


🏝️ Benefits

• 15 days of vacation.

• Floating and company holidays.

• Wellness benefits.

• Paid parental leave.

• Comprehensive medical insurance.

• Dental insurance.

• Vision insurance.

• Life insurance.

• Disability insurance.

• 401(k) plan.

• Work from home stipend.

• Access to a learning platform, training, and tools.

• Flexible work model combining in-person and remote work.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers