
AI Data Engineer III
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in Brazil.
• Develop the knowledge layer for Rimini Street’s Agentic ERP Platform.
• Design and construct RAG pipelines to fetch relevant context for AI agent responses.
• Implement chunking, hybrid retrieval, query reformulation, retrieval evaluation, and reranking pipelines.
• Manage and implement vector storage utilizing PostgreSQL and pgvector.
• Assess embedding models and create scalable embedding pipelines.
• Execute incremental indexing and multi-tenant vector architectures.
• Monitor and enhance vector search latency, accuracy, and resource efficiency.
• Ingest knowledge from Salesforce, ServiceNow, documentation repositories, email archives, and ERP transaction logs.
• Create ETL processes to cleanse, normalize, and enrich raw data.
• Develop pipelines for PDF extraction, HTML parsing, and structured data normalization.
• Create connectors for Salesforce, ServiceNow, SharePoint, and Confluence.
• Implement data quality monitoring, alerting, and lineage tracking.
• Design knowledge architecture following the Four-Spoke model.
• Construct knowledge graphs, relationship models, and metadata taxonomies.
• Create knowledge versioning strategies and feedback loops for ongoing enhancement.
• Integrate with Snowflake and create pipelines between operational systems, Snowflake, and vector stores.
• Implement secure data access patterns to ensure customer isolation.
• Design synchronization between the cloud data warehouse and real-time retrieval systems.
• Optimize query patterns for economically efficient processing at scale.
• Report to the Sr. Director of Engineering.
• Collaborate with GenAI Engineers and articulate data architecture decisions to both technical and non-technical stakeholders.
• Over 5 years of experience in data engineering.
• A minimum of 1–2 years dedicated to AI/ML data pipelines or RAG systems.
• Practical experience in building and enhancing RAG pipelines within production settings.
• Extensive experience with vector databases, embeddings, and similarity search.
• Proficient in ETL/ELT pipelines and data integration from various source systems.
• Production experience with PostgreSQL and SQL-based data handling.
• Proficient in Python for data processing and pipeline creation.
• Skilled in Python for data engineering: pandas, data processing pipelines, and asynchronous programming.
• Strong SQL capabilities with PostgreSQL, including JSONB, full-text search, and extensions.
• Familiarity with vector databases and embeddings, such as pgvector or similar technologies.
• Understanding of RAG concepts including chunking strategies, embedding models, retrieval techniques, and reranking.
• Experience with ETL/ELT methodologies and orchestration tools like Airflow, Dagster, or Prefect.
• Data modeling skills for both relational and document-oriented scenarios.
• Proficient in Git version control and CI/CD practices for data pipelines.
• Knowledge of REST or GraphQL API integration for data extraction.
• Fluent in English, both written and spoken.
• Bachelor’s or Master’s degree in Computer Science, Data Science, or a related field is preferred.
• Experience within enterprise software companies or B2B SaaS platforms.
• Background in information retrieval, search systems, or NLP.
• Certifications in Snowflake, AWS Data Engineering, or similar fields.
• Contributions to open-source data or AI/ML projects.
• Competitive compensation, bonuses, and benefits that align with the skills of top-performing team members.
• Minimal travel; occasional travel may be required for team meetings or training sessions.
• Flexible remote work environment.
• Commitment to a diverse and inclusive workplace.
• Equal Employment Opportunity.
• Involvement in Rimini Street Foundation's community and philanthropic initiatives.
Mirantis
Get handpicked remote jobs straight to your inbox weekly.