Senior Serverless Spark Migration Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in Brazil, +3 more countries.

πŸ“‹ Description

β€’ Lead the migration of enterprise Spark workloads from on-premise settings to AWS and GCP.

β€’ Evaluate Spark applications, clusters, configurations, dependencies, data flows, and resource utilization.

β€’ Identify migration strategies, including rehosting, replatforming, refactoring, modernization, or retirement.

β€’ Upgrade traditional cluster-based workloads to serverless Spark where applicable.

β€’ Design and implement architectures utilizing AWS EMR Serverless, S3, Glue, Lake Formation, GCP Dataproc Serverless, GCS, and BigQuery.

β€’ Refactor legacy PySpark/Scala/Spark SQL applications for cloud compatibility, scalability, and reliability.

β€’ Migrate workloads using Hadoop, HDFS, YARN, Hive, and on-premise Spark clusters.

β€’ Diagnose and enhance Spark workloads, focusing on partitioning, shuffle behavior, joins, data skew, execution plans, executor configuration, serialization, and SQL execution.

β€’ Conduct performance benchmarking and optimize serverless workloads for efficiency, reliability, and cloud cost management.

β€’ Develop reusable migration tools, automation, templates, and frameworks.

β€’ Implement CI/CD and Infrastructure as Code using tools like Terraform.

β€’ Define testing, validation, cutover, rollback, observability, and production-readiness protocols.

β€’ Collaborate with teams in Data Engineering, ML, Cloud Architecture, Platform Engineering, DevOps/SRE, Security, Governance, and FinOps.

β€’ Oversee the entire migration lifecycle: Discover, Assess, Design, Refactor, Migrate, Validate, Optimize, and Operate.


⛳️ Requirements

β€’ Over 8 years of experience in data engineering, distributed systems, cloud engineering, or platform engineering.

β€’ More than 5 years of hands-on experience with Apache Spark in enterprise environments.

β€’ Strong development experience in PySpark and/or Scala.

β€’ Demonstrated experience in migrating large-scale Spark workloads between different infrastructure platforms.

β€’ Practical experience with both AWS and GCP.

β€’ Familiarity with on-premise Hadoop/Spark ecosystems, including HDFS, YARN, and Hive.

β€’ In-depth understanding of Spark internals and distributed processing.

β€’ Solid foundation in SQL and data engineering principles.

β€’ Experience with cloud data lakes and object storage solutions.

β€’ Strong skills in production troubleshooting and performance optimization.

β€’ Proficiency with CI/CD, Git, and Infrastructure as Code.

β€’ Capability to manage migration projects end-to-end, from discovery and architecture to production cutover and optimization.

β€’ Experience with EMR/EMR Serverless, Dataproc/Dataproc Serverless, Glue, Lake Formation, BigQuery, Delta Lake, Iceberg, Kafka, Airflow, Terraform, Docker, or Kubernetes is advantageous.


🏝️ Benefits

β€’ Flexible remote work arrangement.

β€’ Coverage during Pacific Hours (8:00 AM–5:00 PM PST).

People also viewed

Mercor13 hours ago

Industrial Engineer

US flagUnited States OnlyFreelanceEngineer$70 – $110/hour
ApplyView job
RTX22 hours ago

Principal Equipment Engineer, Adhesion & Assembly

US flagFlorida OnlyFull-timeEngineer$107.5k – $204.5k/year
ApplyView job
Expel23 hours ago

Senior Detection & Response Engineer

US flagVirginia OnlyFull-timeEngineer$142.9k – $207.2k/year
ApplyView job
Qualus1 day ago

Engineer II – Relay Settings, System Protection

US flagFlorida, +2 more statesFull-timeEngineer
ApplyView job
General Dynamics Information Technology1 day ago

Senior VDI / Citrix Engineer

US flagUnited States OnlyFull-timeEngineer$149.5k – $195.5k/year
ApplyView job
ARCOS LLC1 day ago

Configuration Engineer

PL flagPoland, +4 more countriesFreelanceEngineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers