
Data Engineer
Posted 5 hours ago

Posted 5 hours ago
This is a fully remote position, open to applicants in Poland.
• Design, construct, and sustain batch and streaming data pipelines using Python, including orchestration, scheduling, and monitoring via Airflow.
• Develop reusable, well-typed Python libraries and internal packages utilized by engineers and analysts.
• Create Python services and APIs with FastAPI and integrate them with both third-party and internal APIs.
• Implement SQL and dbt data transformations on Snowflake and establish data models for efficient and reliable querying.
• Write unit, integration, and data-contract tests while ensuring automated CI coverage is maintained.
• Profile and enhance Python code and data-processing jobs to optimize runtime, memory usage, and costs.
• Deploy and manage code in the cloud using containers, infrastructure as code, and CI/CD methodologies.
• Enforce access controls, manage secrets, and protect sensitive data throughout the data lifecycle.
• Effectively utilize AI coding assistants, ensuring generated outputs are validated before production deployment.
• Identify and resolve issues related to efficiency, reliability, cost, and correctness in existing data pipelines.
• Substitute one-off scripts and notebooks with thoroughly tested, packaged, and scheduled code.
• Implement data quality checks, validation, and monitoring processes.
• Develop matching logic to deduplicate and link entities across various data sources.
• Document data processes and system architecture while maintaining comprehensive project documentation.
• Proficient in Python: skilled in data structures, typing, error handling, generators/iterators, context managers, and the standard library.
• Strong foundation in software engineering principles: modular design, dependency management, packaging, and API design.
• Committed to testing practices: familiar with pytest, fixtures, mocking, and creating code that is inherently testable.
• Practical experience in building and maintaining ETL/ELT pipelines in a production environment rather than just scripts or notebooks.
• Extensive experience with Snowflake and dbt.
• Familiarity with Apache Airflow or similar code-based orchestration tools.
• Solid understanding of SQL and data modeling, capable of designing robust database schemas and executing effective queries.
• Experience with Docker, Kubernetes, and CI/CD methodologies.
• Debugging and profiling abilities; capable of analyzing performance, concurrency, and memory in Python applications.
• Experience with at least one public cloud platform (AWS or Azure).
• Proficient in version control systems (Git).
• Able to articulate technical trade-offs to both technical and non-technical stakeholders.
• Regular use of AI coding assistants such as Claude Code, Cursor, or similar tools.
• Strong communication abilities and a good command of English (minimum C1 level).
• Familiarity with Apache Spark, preferably on Databricks.
• Experience with Pydantic or similar libraries for schema validation.
• Knowledge of Python web/API frameworks (FastAPI, Flask).
• Understanding of Async Python, multiprocessing, or other concurrency patterns.
• Experience with Azure AI Search or AWS OpenSearch.
• Proficiency in a second programming language: Go, Rust, Scala, or TypeScript.
• Familiarity with LLMs, Azure OpenAI, or agentic AI systems.
• Flexible working hours and options for remote, office, or hybrid work arrangements.
• Opportunities for professional growth facilitated by internal training sessions and a dedicated training budget.
• Comprehensive onboarding process with a hands-on approach to ensure a smooth start.
• A supportive atmosphere among professionals who are passionate about their work.
• The possibility to switch projects that align with your interests.
Expleo Group
M3 USA
M3 USA
Providence
Get handpicked remote jobs straight to your inbox weekly.