Remotery

Lead DevOps/AIOps Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in Maryland.

📋 Description

• Oversee the development and execution of cloud-native DevOps and MLOps frameworks on Google Cloud Platform (GCP).

• Create and enhance CI/CD pipelines for data, machine learning, and application workloads.

• Construct infrastructure-as-code utilizing tools like Terraform, establishing repeatable deployment methodologies.

• Design and operationalize data platforms leveraging BigQuery, Cloud Storage, Dataflow, Pub/Sub, Dataproc, and Cloud Composer.

• Develop MLOps functionalities encompassing model creation, deployment, monitoring, version management, and retraining.

• Set up observability across data and ML platforms, which includes logging, monitoring, alerting, pipeline health, data quality, and model performance assessment.

• Implement secure, scalable cloud infrastructure using GCP IAM, networking solutions, secrets management, and security protocols.

• Collaborate with Data Engineers, ML Engineers, Architects, and client stakeholders to convert business needs into production-ready systems.

• Define engineering standards for automation of deployments, testing, environment management, reliability, and operational excellence.

• Address intricate production challenges, leading root-cause analysis and long-term solutions.

• Guide engineers and act as a technical authority across DevOps, cloud, data, and MLOps projects.

• Assess emerging GCP and AI technologies for their business and engineering implications.


⛳️ Requirements

• Over 7 years of experience in DevOps, cloud engineering, platform engineering, MLOps, or a related field.

• Robust hands-on experience with Google Cloud Platform, especially with BigQuery and cloud-native data services.

• Proven experience in designing and executing comprehensive data platforms on GCP.

• Deep understanding of BigQuery architecture, performance enhancement, data ingestion, partitioning, clustering, and data security.

• Familiarity with CI/CD, Git, automated testing, containerization, and Kubernetes/GKE.

• Strong knowledge of Infrastructure-as-Code practices, particularly with Terraform.

• Experience with Vertex AI and/or operational ML platforms, including model deployment and monitoring.

• Proficient in Cloud Composer/Airflow, Dataflow, Dataproc/Spark, and Pub/Sub.

• Solid grasp of observability, reliability engineering, monitoring, logging, and alerting practices.

• Proficient in Python and/or Bash scripting.

• Strong understanding of cloud security, IAM, networking, secrets management, and enterprise governance.

• Capable of functioning at both architectural and practical engineering levels.

• Excellent communication abilities and adeptness at collaborating with technical teams and senior client stakeholders.

• Nice to have: Experience with Vertex AI, MLflow, Kubeflow, GenAI/LLM production solutions, as well as Docker and Kubernetes/GKE, data quality, data lineage, metadata management, semantic data layers, AWS or Azure, and consulting/professional services experience.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Flexible work hours and remote work options.

• Opportunities for professional development and continuous learning.

• Comprehensive health and wellness benefits.

• Engaging company culture with team-building activities.

People also viewed

SYNCREON5 hours ago

Forward Deployment Engineer – Travel to Boston, MA as required

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Rimutee6 hours ago

DevOps, AWS

US flagUnited States OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Mirantis6 hours ago

Senior Site Reliability Engineer, Golang, Kubernetes

CA flagCanada OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group6 hours ago

DevOps Engineer

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
XTEL6 hours ago

DevOps Engineer

BE flagBelgium OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Hypoport SE6 hours ago

DevOps Engineer – m/f/d

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers