
Senior Data Engineer
Posted Aug 20

Posted Aug 20
This is a fully remote position, open to applicants in Poland.
• Design and create scalable, cloud-native data platforms from inception to production.
• Implement near-real-time data ingestion pipelines utilizing event-driven patterns.
• Establish and uphold platform standards, which include Data Lake / Lakehouse principles, medallion architecture, and data contracts.
• Refactor and enhance existing Spark and PySpark scripts for improved performance and maintainability.
• Introduce best practices for code quality, testing, and CI/CD within data pipelines.
• Promote the use of AI tools and agentic workflows within the data engineering team.
• Ensure data quality, observability, and reliability throughout all pipelines and platforms.
• Create self-service tools and microservices to facilitate platform use for other teams.
• Collaborate with Machine Learning, Data Science, and Product teams.
• Lead greenfield projects, cloud migrations, and research & development surrounding agentic AI architectures, event-driven systems, and LLM-ready data pipelines.
• Over 5 years of professional experience in Data Engineering.
• Strong development skills in Python and SQL for pipeline development and optimization.
• Proficient in Apache Spark / PySpark, including query optimization and performance tuning.
• Hands-on experience with Databricks (preferred) or Snowflake.
• Familiarity with at least one major cloud provider: Azure (preferred), AWS, or GCP.
• Experience with stream processing technologies such as Kafka and Spark Structured Streaming.
• Solid understanding of ETL/ELT patterns, data modeling (dimensional, Data Vault), and data warehousing.
• Experience with orchestration tools like Apache Airflow, Azure Data Factory, or equivalent.
• Knowledge of Infrastructure as Code (Terraform or equivalent).
• Understanding of the requirements for production-grade systems: reliability, scalability, observability, and performance.
• Upper-Intermediate level of English proficiency.
• Familiarity with RAG pipeline design and LLM integration patterns.
• Knowledge of data governance frameworks and tools such as Unity Catalog, Apache Atlas, or similar.
• Experience using dbt for data transformation and modeling.
• Familiarity with MLflow, Feature Stores, or ML platform integration.
• Self-motivated and proactive in identifying areas for improvement.
• Comfortable working in a dynamic, fast-paced environment.
• Strong problem-solving skills with a keen attention to detail.
• Open to experimenting with emerging technologies and methodologies.
• Option for remote work.
• Full-time employment.
Expleo Group
CodiLime
M3 USA
M3 USA
Get handpicked remote jobs straight to your inbox weekly.