
Cloud Engineer – Azure Databricks
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Mexico.
• Design, implement, and manage the Azure infrastructure foundation for an enterprise data and AI platform.
• Create and maintain reusable Terraform modules for provisioning Databricks workspaces, ADLS Gen2, Data Factory, Key Vault, networking, and Azure ML.
• Oversee remote state management, module versioning, drift detection, and infrastructure deployment pipelines.
• Develop private networking and security architecture, including private endpoints, hub-and-spoke topology, VNet injection, NSGs, firewall rules, managed identities, RBAC, and data exfiltration controls.
• Manage ADLS Gen2 storage architecture, access control lists (ACLs), lifecycle and tiering policies, encryption, key management, and access patterns.
• Operate Azure Data Factory integration runtimes, linked-service credentials, managed identities, environment promotion, deployment automation, and monitoring.
• Provision and manage Azure ML workspaces, compute clusters, GPU capacity, MLflow model registry integration, model-serving endpoints, identity, and networking.
• Build and maintain observability using Azure Monitor, Log Analytics, diagnostic settings, alerting, and operational runbooks.
• Take ownership of FinOps visibility and optimization for DBU and storage expenditures, including tagging, chargeback/showback, budget alerts, reserved capacity, and capacity planning.
• Manage CI/CD pipelines and promotion across development, testing, and production environments.
• Establish standards for Databricks account and workspace administration, topology, settings, role delegation, and multi-workspace strategy.
• Design Unity Catalog metastores, permissions, storage credentials, external locations, lineage, audit, and Delta Sharing.
• Manage Entra ID integration, SCIM, identity federation, groups, entitlements, service principals, and token policies.
• Define compute governance through cluster policies, instance pools, node/runtime standards, autoscaling, autotermination, Photon/serverless evaluation, and SQL warehouse configuration.
• Analyze system tables, usage attribution, and DBU forecasting to identify spending concentrations and remediation paths.
• Review platform configurations, guide architectural decisions, publish standards and self-service patterns, and support administration and data teams.
• Diagnose job failures, cluster startup issues, permission errors, connectivity faults, and performance concerns across infrastructure, Databricks configuration, and workloads.
• Engage in production ETL/ELT development, Spark transformations, model development, dimensional modeling, dbt, and BI/semantic-layer collaboration with other teams.
• 6+ years of experience in cloud infrastructure or platform engineering, with a minimum of 4 years dedicated to Microsoft Azure.
• Expert-level, hands-on experience with Terraform, including production modules, remote state management, and CI/CD infrastructure as code.
• Proven hands-on experience in administering Databricks: account and workspace administration, Unity Catalog, cluster policies, identity federation, and cost governance.
• A history of serving as a senior technical resource to adjacent teams, guiding stakeholders and fostering consensus.
• In-depth knowledge of Azure networking and security, including private endpoints, VNets, hub-and-spoke design, NSGs, Entra ID, managed identities, RBAC, and Key Vault.
• Experience in enterprise-scale ADLS Gen2 design and access control.
• Familiarity with Azure Data Factory platform operations, including integration runtimes, managed VNet, credential management, and deployment automation.
• Proficient in Python, as well as PowerShell and/or Bash for platform tooling and automation.
• Basic knowledge of Spark and SQL for diagnosing performance issues at the infrastructure and configuration levels.
• Preference for experience with Databricks Asset Bundles, Terraform Databricks provider, and workspace-as-code patterns.
• Experience with Unity Catalog migration or multi-workspace consolidation is preferred.
• Familiarity with Azure Machine Learning, MLflow, or model-serving infrastructure is a plus.
• Preferred experience with Kubernetes/AKS and containerized workloads.
• Familiarity with policy as code, event-driven infrastructure, formal FinOps practices, and multi-region or multi-tenant Databricks experience is desired.
• Preferred certifications include Databricks platform credentials, Azure certifications, or HashiCorp Terraform Associate.
• Advanced English proficiency is required.
• Remote work opportunities.
• Fully remote position.
• Opportunity to practice advanced English in a remote client environment.
inexogy smart metering
Skylo
SYNCREON
TechPionier
Get handpicked remote jobs straight to your inbox weekly.