
Staff Data Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Colombia, +1 more country.
• Define the architecture and strategy for data platforms, leading the design of pipelines, warehouses, and data lakes.
• Construct and enhance scalable data pipelines that facilitate both batch and real-time processing.
• Establish and uphold data governance, quality standards, and compliance frameworks throughout the platform.
• Develop monitoring, logging, and alerting systems for data pipelines and services, while contributing to CI/CD workflows for data deployment and automation.
• Propel the modernization of the data platform, focusing on performance, cost efficiency, and scalability.
• Utilize tools such as Claude, Cursor, and other contemporary AI assistants to deliver high-quality work efficiently.
• Design and execute data contracts and event flows in collaboration with backend, platform, and engineering teams.
• Lead the design and execution of data pipelines for production AI/ML systems, encompassing embeddings, vector stores, RAG data preparation, feature stores, and training/inference data flows.
• Integrate data services with APIs, middleware, and external systems.
• Collaborate with leadership on the data strategy.
• Work together with engineering, analytics, AI, and product teams to align data platforms with overarching objectives.
• Promote data quality, governance, and platform best practices across various teams and projects.
• Set data engineering standards within the team.
• Mentor junior and mid-level engineers.
• Make significant architectural decisions with clear accountability and consideration of long-term implications.
• Over 7 years of professional experience in data engineering, with a strong background in leading intricate data platform projects.
• Robust system architecture experience with a focus on distributed data systems.
• Expert-level proficiency in Python, Scala, and SQL.
• In-depth knowledge of cloud-native data platforms and enterprise data warehousing.
• Strong expertise in orchestrating and processing data pipelines.
• Extensive experience with streaming platforms and real-time data processing (e.g., Kafka, Kinesis, Pub/Sub).
• Profound data modeling skills along with experience in data transformation.
• Significant experience with data quality, governance, and compliance frameworks.
• Considerable experience in container orchestration and CI/CD practices for data systems.
• Proven experience in building data pipelines for production AI/ML systems, including embeddings, vector stores, RAG data preparation, feature stores, and training/inference data flows.
• Demonstrated leadership and technical mentoring capabilities within a team or organization.
• Excellent communication skills with the ability to convey technical concepts to diverse audiences.
• Proven day-to-day utilization and expert knowledge of AI-forward coding tools such as Claude and Cursor.
• Strong problem-solving abilities with a knack for handling complex technical and business challenges using sound judgment.
• Experience with data mesh or data fabric concepts, lakehouse architectures, or the implementation of governance frameworks is advantageous.
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Opportunities for professional development and career growth.
• Comprehensive health and wellness benefits.
• Collaborative and innovative work environment.
Magna Legal Services
Huron
Strategic Systems International
Get handpicked remote jobs straight to your inbox weekly.