
Lead DevOps/AIOps Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Maryland.
• Oversee the development and execution of cloud-native DevOps and MLOps frameworks on Google Cloud Platform (GCP).
• Create and enhance CI/CD pipelines for data, machine learning, and application workloads.
• Construct infrastructure-as-code utilizing tools like Terraform, establishing repeatable deployment methodologies.
• Design and operationalize data platforms leveraging BigQuery, Cloud Storage, Dataflow, Pub/Sub, Dataproc, and Cloud Composer.
• Develop MLOps functionalities encompassing model creation, deployment, monitoring, version management, and retraining.
• Set up observability across data and ML platforms, which includes logging, monitoring, alerting, pipeline health, data quality, and model performance assessment.
• Implement secure, scalable cloud infrastructure using GCP IAM, networking solutions, secrets management, and security protocols.
• Collaborate with Data Engineers, ML Engineers, Architects, and client stakeholders to convert business needs into production-ready systems.
• Define engineering standards for automation of deployments, testing, environment management, reliability, and operational excellence.
• Address intricate production challenges, leading root-cause analysis and long-term solutions.
• Guide engineers and act as a technical authority across DevOps, cloud, data, and MLOps projects.
• Assess emerging GCP and AI technologies for their business and engineering implications.
• Over 7 years of experience in DevOps, cloud engineering, platform engineering, MLOps, or a related field.
• Robust hands-on experience with Google Cloud Platform, especially with BigQuery and cloud-native data services.
• Proven experience in designing and executing comprehensive data platforms on GCP.
• Deep understanding of BigQuery architecture, performance enhancement, data ingestion, partitioning, clustering, and data security.
• Familiarity with CI/CD, Git, automated testing, containerization, and Kubernetes/GKE.
• Strong knowledge of Infrastructure-as-Code practices, particularly with Terraform.
• Experience with Vertex AI and/or operational ML platforms, including model deployment and monitoring.
• Proficient in Cloud Composer/Airflow, Dataflow, Dataproc/Spark, and Pub/Sub.
• Solid grasp of observability, reliability engineering, monitoring, logging, and alerting practices.
• Proficient in Python and/or Bash scripting.
• Strong understanding of cloud security, IAM, networking, secrets management, and enterprise governance.
• Capable of functioning at both architectural and practical engineering levels.
• Excellent communication abilities and adeptness at collaborating with technical teams and senior client stakeholders.
• Nice to have: Experience with Vertex AI, MLflow, Kubeflow, GenAI/LLM production solutions, as well as Docker and Kubernetes/GKE, data quality, data lineage, metadata management, semantic data layers, AWS or Azure, and consulting/professional services experience.
• Competitive salary and performance-based bonuses.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Comprehensive health and wellness benefits.
• Engaging company culture with team-building activities.
SYNCREON
Rimutee
Mirantis
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.