
RAG Architect
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in India.
• Design and implement comprehensive Retrieval-Augmented Generation (RAG) pipelines for document analysis, semantic metadata enhancement, and multi-vector searches.
• Enhance context window efficiency through parent-child chunking strategies, sentence-window retrieval methods, and sliding window approaches.
• Develop high-performance re-ranking components utilizing machine learning cross-encoders such as Cohere Rerank and BGE-Reranker.
• Create automated semantic caching frameworks leveraging caching solutions like GPTCache.
• Set up automated data chunking systems for PDFs, corporate wikis, and SQL outputs.
• Manage vector similarity spaces by refining hybrid search algorithms that integrate dense semantic embeddings with sparse keyword token indexes like BM25.
• Evaluate context-level hallucination rates and accuracy logs, monitoring precision metrics, retrieval recall thresholds, and processing speeds.
• Take ownership of enterprise generative AI retrieval performance, accuracy, and operational cost metrics.
• 6 to 10 years of experience in enterprise data engineering, database architecture, or search engine development.
• A minimum of 3 years dedicated to scaling context retrieval loops for live large language model (LLM) applications.
• Required certification: Google Cloud Certified Professional Cloud Database Engineer, AWS Certified Data Analytics - Specialty, Databricks Certified Data Engineer Professional, or an equivalent Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud provider (AWS/GCP/Azure).
• Strong expertise in Python programming.
• Proficient in vector databases, including Pinecone, Milvus, and Weaviate.
• Advanced knowledge of text embedding models.
• Strong command of open-source orchestration tools such as LlamaIndex and LangChain.
• Solid understanding of SQL.
• In-depth comprehension of context window constraints, particularly the “lost in the middle” phenomenon.
• Familiarity with multi-modal token dynamics.
• Insight into network data transfer velocities.
• Awareness of cloud memory architectures.
• Previous experience in implementing Graph RAG frameworks using native knowledge graphs like Neo4j is preferred.
• Experience in fine-tuning open-source text embedding models for industry-specific terminology or legacy product schemas is preferred.
• Flexible remote work options.
• Contract-based employment.
Fortive
Strada
Salesforce
FyerX - Your Trusted Marketing Partner
Get handpicked remote jobs straight to your inbox weekly.