
Lead Data Platform Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in New York.
• Oversee the complete **Data pipeline** (ETL jobs) ensuring compliance with established SLAs.
• Administer AWS core and **big data services** (S3, IAM, EMR, Redshift, etc.).
• Execute applications within container environments (ECS, Docker).
• Spearhead the Day 2 operational lifecycle for ML and GenAI infrastructure, which involves designing, deploying, and maintaining high-availability production platforms for LLM serving. Implement automated scaling, self-healing, and infrastructure-as-code methodologies, with an emphasis on proactive reliability, model performance observability, and continuous cost optimization for high-compute AI workloads.
• Work in close collaboration with product development and engineering teams to develop AI-driven features.
• Enhance cloud operations consistency by automating platform maintenance, standardizing infrastructure configurations (IaC), and establishing robust release management processes to minimize drift across multi-cloud environments.
• Manage AWS infrastructure through code (Terraform, Chef, etc.).
• Administer applications operating on Linux systems.
• Facilitate application and system monitoring to improve observability.
• Provide application and infrastructure support for ETL jobs and data pipelines, including participation in an on-call rotation for after-hours emergencies.
• Collaborate with platform and development teams to strategize and deploy product releases and patch Linux/ECS clusters.
• Demonstrate the ability to engage in design reviews, code reviews, and troubleshoot incidents.
• Capable of functioning in high-pressure environments and quickly resolving complex issues while effectively managing multiple priorities.
• Proficient in documenting, writing, and reviewing RCAs.
• Bachelor's Degree with a minimum of 8+ years of experience in managing Big Data technologies and Data Pipelines.
• Strong knowledge and experience in Linux administration and troubleshooting.
• Over 5 years of experience managing cloud infrastructure and platforms, including AWS and Azure.
• Familiarity with the current engineering landscape in generative AI, coupled with a strong passion for AI and related technologies.
• Expertise in MLOps and operations for production-grade LLMs. Proven experience in managing high-availability model inference clusters, automating model lifecycle management, and implementing advanced observability (latency, throughput, and error rate monitoring) specifically for AI workloads.
• Proficient in Bash or Python scripting.
• Experience with containerization technologies, including Amazon ECS, EKS/Azure AKS.
• Familiarity with tools like Chef, Ansible, Jenkins, Rundeck, or similar.
• Proficient in source control systems such as Git and adept at navigating complex branching strategies.
• Experience with Infrastructure as Code tools such as Terraform and helm charts.
• Solid understanding of DNS and load balancer setup and troubleshooting.
• Experience in Big Data platforms/Data lakes and managing Business Intelligence tools (like Looker).
• Knowledge of Apache Spark architecture and troubleshooting Java applications.
• Basic understanding of MySQL Server and general database concepts.
• Excellent written and verbal communication skills with a strong inclination for problem-solving.
• Confidence in your ability to independently manage and deliver projects and resolve issues, with a global perspective.
• Extensive experience in Day 2 cloud operations, including automated incident remediation, capacity planning, and managing large-scale production cloud environments with a focus on performance and reliability.
• Opportunity to work with cutting-edge AI technologies.
• Collaborative and innovative work environment.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional growth and development.
Arctiq
Cisco
Prove
Hello Heart
Get handpicked remote jobs straight to your inbox weekly.