
Architect
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in Mexico.
• Develop and execute comprehensive data architectures utilizing the Databricks Lakehouse Platform along with Medallion Architecture.
• Supervise the consistent generation of production-ready data pipelines and organization methodologies.
• Diagnose and optimize large-scale distributed computing tasks by leveraging Apache Spark internals.
• Oversee data schema management, ACID transactions, and optimization processes in Delta Lake, including Z-ordering.
• Create holistic governance frameworks, encompassing row/column-level security, RBAC, and data lineage through Unity Catalog.
• Set up native cloud infrastructure for Azure, AWS, or GCP, which includes VNet configurations and IAM roles.
• Track resource utilization and DBUs, appropriately size serverless or multi-node clusters, and develop strategies for cost optimization.
• Extensive knowledge of the Databricks Lakehouse Platform, Apache Spark, and Delta Lake.
• Proven experience in designing and implementing Medallion data architectures.
• Strong understanding of Spark internals and large-scale performance optimization.
• Demonstrated proficiency with Unity Catalog for data governance and management of AI assets.
• Solid comprehension of native cloud environments (AWS, Azure, or GCP).
• Advanced programming skills in SQL and either Python or Scala.
• Proven track record in cloud cost optimization and mitigating unexpected expenses.
• Resilience, emotional intelligence, and a focus on agile delivery methodologies.
• Profound insight into the distinction between coding and engineering practices.
• Proficient in oral English.
• Proficient in Spanish.
• Valid Databricks Certified Data Architect or Databricks Certified Data Engineer Professional certifications (preferred).
• Experience in building and deploying GenAI applications and ML models using Databricks Model Serving, Vector Search, and the Mosaic AI suite (preferred).
• Familiarity with automating Databricks workflows through Git, Databricks Repos, dbt, Azure Data Factory, or Apache Airflow (preferred).
• Understanding of Databricks Structured Streaming and Auto Loader (preferred).
• Knowledge of Databricks SQL, Serverless Warehouses, and BI tool integration such as Power BI (preferred).
• Expertise in designing end-to-end data architectures across Bronze, Silver, and Gold layers.
• In-depth understanding of Databricks and Spark internals, including troubleshooting, tuning, caching, partitioning, and broadcast joins.
• Knowledge of ACID transactions, schema enforcement, time travel, and optimization practices in Delta Lake.
• Proven capability in designing unified governance frameworks using Unity Catalog, including RBAC, row/column-level security, and data lineage.
• Advanced SQL capabilities and fluency in Python or Scala.
• Ability to monitor DBUs and execute cost optimization and performance tuning strategies.
• Flexible remote work arrangements.
• Opportunities for global career advancement and involvement in impactful projects.
• Engage in comprehensive digital transformation initiatives for leading global organizations.
• Collaborate within an international network of experts.
Applaudo
Novartis
Salesforce
MWDN
Get handpicked remote jobs straight to your inbox weekly.