
Senior Data Platform Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Design, develop, and sustain scalable data and machine learning pipelines utilizing Python and SQL across both batch and streaming workloads.
• Take ownership of and enhance the Databricks platform infrastructure, which includes Delta Lake architecture, governance via Unity Catalog, orchestration through Databricks Workflows, and compute optimization.
• Create and manage end-to-end machine learning pipelines that encompass feature engineering, model training, experiment tracking, and model deployment/serving.
• Collaborate with data scientists to bring models into production-grade machine learning systems.
• Define and apply standards for the data platform, including ingestion patterns, data modeling conventions, medallion architecture, and practices for reliability.
• Implement frameworks for data quality, observability, and monitoring.
• Optimize pipelines for performance, cost efficiency, and reliability using Spark and PySpark.
• Assess, integrate, and govern platform tools and data sources within the Databricks ecosystem.
• Contribute to architectural decisions and the strategic roadmap for the data platform.
• Engage in code reviews, technical design discussions, and establish engineering standards.
• Mentor junior engineers and enhance platform and data engineering practices.
• Document the architecture of the platform, design of pipelines, and operational runbooks.
• Over 6 years of experience in data engineering, data platform development, or machine learning engineering roles.
• Strong command of Python and SQL with hands-on experience in production-grade data pipelines.
• Practical expertise with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, and PySpark.
• Experience in constructing and maintaining production machine learning pipelines, including feature engineering, training, experiment tracking, and model deployment.
• Familiarity with MLflow or other similar experiment tracking and model registry tools.
• Experience with cloud data platforms such as AWS, Azure, or GCP.
• Strong understanding of data modeling, dimensional design, and analytics-friendly data architecture.
• Familiarity with batch and incremental/CDC pipeline patterns.
• Proficient in Git, version control systems, and CI/CD practices for data and ML workflows.
• Strong engineering judgment with a focus on reliability, maintainability, and cost-effectiveness.
• Excellent communication skills and comfort in engaging with both technical and non-technical stakeholders.
• Nice-to-have: experience with streaming or near real-time pipelines, feature stores, LLM/RAG or AI/BI tools, data quality and observability tools, dbt, infrastructure as code, Agile/Scrum methodologies, and mentoring or platform standards experience.
• Generous time-off policies.
• Comprehensive benefits package.
• Support for education, wellness, and lifestyle initiatives.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.