
Senior RAG Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Norway.
• Take ownership of the retrieval pipeline from start to finish: including document parsing and chunking, embedding, indexing, and ensuring all components are up-to-date as clients add or modify material.
• Develop our search strategy by integrating both semantic and keyword retrieval in Qdrant, merging ranked lists, applying metadata filters, and reranking to ensure the few results provided to the model are the most relevant.
• Create agentic retrieval mechanisms: query decomposition, tools for models to search and navigate documents, multi-step loops that determine when to cease operations, while adhering to cost and latency constraints.
• Establish the evaluation framework that verifies the effectiveness of our solutions: including golden sets, retrieval metrics, regression tests with realistic client data, and a tracing system robust enough to identify the chunk responsible for a subpar answer.
• Deliver reliable services rapidly using Python and FastAPI, with intensive tasks such as ingestion, embedding, and reindexing operating as background jobs capable of handling substantial loads.
• Prioritize data security and isolation, ensuring that client-sensitive material is handled with care and that tenant data remains thoroughly segregated.
• Ensure retrieval performance is consistent across various languages used by our clients, matching the quality of English results.
• A minimum of five years of experience in developing backend systems that operate in a production environment, with at least two years dedicated to deploying retrieval or RAG systems relied upon by real users.
• Proven experience diagnosing and resolving retrieval regressions in a production setting. You can articulate what went wrong, how you identified the issue, and the metrics before and after the fix.
• Experience in building and maintaining a golden set: detailing the number of queries, who performed the labeling, the metrics you trust, and a decision you made based on their insights.
• Practical experience running a vector index in production: selecting index parameters, managing the trade-offs between memory and latency, and reindexing without causing downtime in search functionality.
• Proficient in Python. Familiarity with FastAPI or a similar framework, and capable of managing PostgreSQL and Redis under load.
• You do not implement a retrieval change based solely on improved output from a limited set of manual queries.
• You take full responsibility for projects from beginning to end, including routine maintenance and addressing technical debt, instead of focusing solely on the more intriguing aspects of development.
• You stay informed about advancements in vector databases and LLMs out of genuine curiosity, not just as part of your job. Share your side project, a benchmark you completed for enjoyment, or any repository where you experimented before it became common practice.
• Experience collaborating with AI coding assistants, with insights into their usefulness and limitations.
• Comfortable in a startup environment that may pivot direction. Understand that decisions are made and may be revisited.
• Fully remote work opportunity across the EEA, including Norway.
• You must possess the existing right to work in your current location.
• Be a key player in driving innovation within one of the most thrilling AI ventures.
• Collaborate with skilled colleagues in a dynamic, high-energy team.
• Contribute to shaping both the product and culture as we continue to grow.
• Enjoy a flexible, English-speaking work environment.
LiteLLM AI Gateway
Snowflake
RTX
C-MORE
Get handpicked remote jobs straight to your inbox weekly.