
Data and Machine Learning Engineer
Posted Jul 3

Posted Jul 3
This is a fully remote position, open to applicants in Colombia.
• Design, develop, and maintain scalable data pipelines to facilitate the ingestion, transformation, and delivery into centralized feature stores, model-training workflows, and real-time inference services.
• Construct and enhance workflows for extracting, storing, and retrieving semantic representations of unstructured data to support advanced search and retrieval patterns.
• Architect and implement streamlined analytics and dashboard solutions that provide a natural language query experience and insights powered by AI.
• Define and execute processes for managing prompt engineering techniques, orchestration flows, and model fine-tuning routines to enhance conversational interfaces.
• Supervise vector data stores and devise efficient indexing methodologies to support retrieval-augmented generation (RAG) workflows.
• Collaborate with data stakeholders to gather requirements for language-model initiatives and translate them into scalable solutions.
• Develop and maintain thorough documentation for all data processes, workflows, and model deployment routines.
• Demonstrate a willingness to stay updated and learn emerging methodologies in data engineering, MLOps, and LLM operations.
• Over 8 years of experience as a Data Engineer, with at least 2 years concentrated on MLOps.
• Exceptional English communication skills.
• Strong oral and written communication abilities with the BI team and user community.
• Proven experience in employing Python for data engineering tasks, including transformation, advanced data manipulation, and large-scale data processing.
• In-depth knowledge of vector databases and RAG architectures, and their role in driving semantic retrieval workflows.
• Proficient in integrating open-source LLM frameworks into data engineering workflows for comprehensive model training, customization, and scalable inference.
• Familiarity with cloud platforms such as AWS or Azure Machine Learning for managed LLM deployments.
• Practical experience with big data technologies including Apache Spark, Hadoop, and Kafka for distributed processing and real-time data ingestion.
• Experience in designing complex data pipelines that extract data from RDBMS, JSON, API, and flat file sources.
• Proven skills in SQL and PLSQL programming, with advanced proficiency in Business Intelligence and data warehouse methodologies, along with hands-on experience in one or more relational database systems and cloud-based database services like Snowflake/Redshift.
• Understanding of software engineering principles and capability to work on Unix/Linux/Windows operating systems, with experience in Agile methodologies.
• Proficient in version control systems, with experience in managing code repositories, branching, merging, and collaborating within a distributed development environment.
• Interest in business operations and a comprehensive understanding of how robust BI systems enhance corporate profitability by fostering data-driven decision-making and strategic insights.
• Competitive salary and performance-based bonuses.
• Opportunities for professional growth and development.
• Flexible work environment and remote work options.
• Comprehensive health benefits including medical, dental, and vision coverage.
• Retirement savings plan with company matching.
Zeta Global
AvidXchange, Inc.
n Human Resources & Management Systems [ nHRMS ]
Get handpicked remote jobs straight to your inbox weekly.