
Senior Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Poland.
• Design and develop scalable, cloud-native data platforms from inception to production.
• Implement near-real-time data ingestion pipelines utilizing event-driven architectures.
• Establish and uphold platform standards, incorporating Data Lake / Lakehouse principles, medallion architecture, and data contracts.
• Refactor and enhance existing Spark and PySpark scripts for better performance and maintainability.
• Introduce best practices for code quality, testing, and CI/CD processes across data pipelines.
• Promote the use of AI tools and agentic workflows within the data engineering team.
• Ensure data quality, observability, and reliability across all data pipelines and platforms.
• Create self-service tools and microservices to facilitate platform usage for other teams.
• Collaborate with Machine Learning, Data Science, and Product teams as a key technical contributor and thought leader.
• Lead research and development initiatives on agentic AI architectures, event-driven systems, and LLM-ready data pipelines, transforming architectural concepts into production-ready solutions.
• Build contemporary cloud-native data platforms, transition on-premises legacy systems to the cloud, and develop AI-ready data infrastructure.
• Over 5 years of professional experience in Data Engineering.
• Strong development skills in Python and SQL for pipeline creation and optimization.
• Proficiency in Apache Spark / PySpark, focusing on query optimization and performance enhancements.
• Practical experience with Databricks (preferred) or Snowflake.
• Familiarity with at least one major cloud service provider: Azure (preferred), AWS, or GCP.
• Experience with stream processing technologies such as Kafka and Spark Structured Streaming.
• Solid grasp of ETL/ELT methodologies, data modeling (dimensional, Data Vault), and data warehousing concepts.
• Experience using orchestration tools like Apache Airflow, Azure Data Factory, or their equivalents.
• Knowledge of Infrastructure as Code practices (Terraform or similar).
• Understanding of production-grade system essentials: reliability, scalability, observability, and performance.
• Upper-Intermediate proficiency in English.
• Familiarity with RAG pipeline design and LLM integration techniques.
• Knowledge of data governance frameworks and tools such as Unity Catalog, Apache Atlas, or similar.
• Experience with dbt for data transformation and modeling.
• Familiarity with MLflow, Feature Stores, or ML platform integration.
• Employees have the option to work remotely.
• Full-time employment available.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.