
Data Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
• Develop, implement, and sustain production-grade data pipelines in line with data science needs.
• Design and construct scalable data flows tailored to specific business applications that can be integrated into the product.
• Construct and uphold machine learning data pipelines.
• Create tools and frameworks that assist data scientists and analytics teams in developing and refining innovative solutions.
• Collaborate closely with Product, R&D, Data, and Analytics teams to improve system functionality and performance.
• Educate customer data scientists and engineers on the maintenance and enhancement of data pipelines within the platform.
• Travel domestically and internationally to provide support to customers as needed.
• Establish and nurture strong technical relationships with customers and partners.
• Minimum of 2 years of practical experience with Apache Spark (required).
• Strong proficiency in SQL.
• Experience utilizing version control systems, especially Git.
• Hands-on experience with the Apache Hadoop ecosystem, including Hive, Impala, Hue, HDFS, and Sqoop.
• Proficiency in Python (Pandas).
• Familiarity with PySpark, Scala, Java, or R.
• Demonstrated experience in data transformation, validation, cleansing, and machine learning feature engineering.
• Bachelor's degree or higher in Computer Science, Statistics, Informatics, Information Systems, Engineering, or a related quantitative field.
• Experience in optimizing large-scale data pipelines, architectures, and datasets.
• Strong analytical skills with experience in working with structured and semi-structured data.
• Experience in building processes that facilitate data transformation, metadata management, dependency management, and workload optimization.
• Ability to conduct root cause analysis on data and processes to address business inquiries and pinpoint improvement opportunities.
• Excellent customer-facing and cross-functional collaboration skills.
• Proficient in English and Spanish, both written and spoken.
• Experience working with Linux.
• Experience in building machine learning pipelines.
• Knowledge of Elasticsearch.
• Familiarity with Zeppelin and/or Jupyter.
• Understanding of workflow orchestration tools such as Jenkins or Apache Airflow.
• Experience with microservices and containerization technologies, including Docker and Kubernetes.
• Competitive salary and benefits package.
• Meaningful opportunities for career development.
• Diverse and inclusive workplace culture.
• Continuous learning and growth opportunities.
• Chance to collaborate with a talented global team.
• A challenging and stimulating work environment.
SysMap Solutions
Get handpicked remote jobs straight to your inbox weekly.