
Senior AI Engineer – LLM, RAG, Agent Systems
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Egypt.
• Design and construct production-quality LLM applications and RAG systems.
• Take ownership of retrieval architecture, encompassing chunking, embeddings, vector search, hybrid search, and reranking.
• Develop agentic and tool-calling systems that incorporate permissions, scoping, validation, and guardrails.
• Create natural-language interfaces for enterprise data and structured databases.
• Establish and maintain LLM evaluation frameworks that include test sets, regression suites, grounding, hallucination, and answer-quality assessment.
• Operate within the data platform and engineering stack instead of depending solely on hosted AI APIs.
• Deploy and enhance self-hosted open-weight models utilizing vLLM or equivalent serving infrastructure.
• Improve inference performance, GPU utilization, latency, throughput, and cost efficiency.
• Investigate and apply fine-tuning or model adaptation as necessary.
• Collaborate with data and software engineers to transform AI capabilities into dependable production products.
• Over 5 years of experience in software or data engineering.
• A minimum of 2 years of practical experience in building and deploying production LLM-based systems.
• Proficient in Python engineering.
• Comprehensive understanding of RAG and retrieval architecture, including chunking strategies, embeddings, vector databases/search, hybrid search, reranking, and retrieval evaluation.
• Experience in building LLM agents or tool-calling systems.
• Familiarity with permissions, access control, scoping, validation, and guardrails for AI systems.
• Strong grasp of LLM evaluation, including test datasets, regression testing, grounding, and detection of hallucinations.
• Experience working directly with data platforms, databases, or enterprise data.
• Solid software engineering fundamentals and capability in transitioning systems from prototype to production.
• Experience with self-hosted open-weight models, vLLM or equivalent model-serving infrastructure, GPU resource management, inference optimization, fine-tuning, LoRA, Text-to-SQL, semantic layers, and integrating unstructured documents with structured enterprise data (strongly preferred).
• Competitive salary and performance-based bonuses.
• Opportunities for professional development and career growth.
• Flexible work arrangements and remote work options.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and holiday policy.
Weekday (YC W21)
Teamficient
Omilia - Conversational Intelligence
Slate Auto
Get handpicked remote jobs straight to your inbox weekly.