Senior AI Engineer – LLM, RAG, Agent Systems

Posted 1 day ago

This is a fully remote position, open to applicants in Egypt.

📋 Description

• Design and construct production-quality LLM applications and RAG systems.

• Take ownership of retrieval architecture, encompassing chunking, embeddings, vector search, hybrid search, and reranking.

• Develop agentic and tool-calling systems that incorporate permissions, scoping, validation, and guardrails.

• Create natural-language interfaces for enterprise data and structured databases.

• Establish and maintain LLM evaluation frameworks that include test sets, regression suites, grounding, hallucination, and answer-quality assessment.

• Operate within the data platform and engineering stack instead of depending solely on hosted AI APIs.

• Deploy and enhance self-hosted open-weight models utilizing vLLM or equivalent serving infrastructure.

• Improve inference performance, GPU utilization, latency, throughput, and cost efficiency.

• Investigate and apply fine-tuning or model adaptation as necessary.

• Collaborate with data and software engineers to transform AI capabilities into dependable production products.


⛳️ Requirements

• Over 5 years of experience in software or data engineering.

• A minimum of 2 years of practical experience in building and deploying production LLM-based systems.

• Proficient in Python engineering.

• Comprehensive understanding of RAG and retrieval architecture, including chunking strategies, embeddings, vector databases/search, hybrid search, reranking, and retrieval evaluation.

• Experience in building LLM agents or tool-calling systems.

• Familiarity with permissions, access control, scoping, validation, and guardrails for AI systems.

• Strong grasp of LLM evaluation, including test datasets, regression testing, grounding, and detection of hallucinations.

• Experience working directly with data platforms, databases, or enterprise data.

• Solid software engineering fundamentals and capability in transitioning systems from prototype to production.

• Experience with self-hosted open-weight models, vLLM or equivalent model-serving infrastructure, GPU resource management, inference optimization, fine-tuning, LoRA, Text-to-SQL, semantic layers, and integrating unstructured documents with structured enterprise data (strongly preferred).


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Opportunities for professional development and career growth.

• Flexible work arrangements and remote work options.

• Comprehensive health, dental, and vision insurance.

• Generous paid time off and holiday policy.

People also viewed

Weekday (YC W21)1 day ago

MLOps Engineer, LLM Systems, Serving, GPU Kernels, Profiling

US flagUnited States, +2 more countriesPart-timeLLM Engineer$90 – $120/hour
ApplyView job
Teamficient1 day ago

AI Engineer – Conversational, Voice & Call Intelligence

US flagUnited States OnlyFreelanceLLM Engineer
ApplyView job
Omilia - Conversational Intelligence2 days ago

Software Engineer – Conversational AI Application Developer, Tech Lead

BR flagBrazil OnlyFull-timeLLM Engineer
ApplyView job
Slate Auto4 days ago

VP – Distinguished Engineer, Generative AI Engineering

US flagUnited States OnlyFull-timeLLM Engineer$222.4k – $370.7k/year
ApplyView job
Belva4 days ago

Senior Generative AI Engineer

US flagUnited States OnlyFull-timeLLM Engineer
ApplyView job
Nagarro5 days ago

Senior Staff Engineer – Generative AI, NLP

US flagPennsylvania OnlyFull-timeLLM Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers