
Senior Engineer, AI
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in India.
• Design, develop, and maintain production AI applications from start to finish, covering backend, frontend, and inference services.
• Architect retrieval-augmented generation (RAG) systems utilizing vector databases, embedding models, and chunking strategies optimized for both accuracy and latency.
• Create agentic workflows incorporating tool and function calling, multi-step reasoning, and structured output parsing, prioritizing accuracy and control.
• Draft and refine system prompts, few-shot examples, and prompt chains to enhance output quality.
• Execute function calling, tool usage patterns, and manage structured JSON/XML output handling using cutting-edge and lightweight models from providers such as Anthropic and OpenAI.
• Drive cost optimization through model selection, caching, token budgeting, and request batching on a large scale.
• Construct and uphold evaluation frameworks to assess accuracy, relevance, hallucination rates, and regression in response to prompt and model changes. Experience with observability tools (Sentry, Opik, etc.) is essential.
• Collaborate with message queues (RabbitMQ), caching layers (Redis), and relational databases (PostgreSQL) that support AI service backends.
• Deploy and manage AI services on Kubernetes using CI/CD pipelines on AWS/GCP.
• Integrate AI functionalities with third-party platforms (such as Telegram bots, chat widgets, etc.).
• Participate in architectural decision-making regarding model selection, hosting (cloud APIs vs. self-hosted), and evaluating build-vs-buy options.
• Over 5 years of experience delivering production software systems.
• At least 2 years of experience in creating AI/LLM-powered applications end-to-end, with real user engagement and volume, not just prototypes.
• Strong expertise in RAG architectures: vector databases, embedding models, chunking/indexing strategies, and retrieval evaluation.
• Comprehensive understanding of LLM capabilities and limitations, including prompt engineering, function/tool calling, structured outputs, context window management, and multi-turn conversations.
• Experience with LLM provider APIs and abstraction layers (OpenAI, Anthropic, LiteLLM, OpenRouter, or similar).
• Proficient in Python (Flask/FastAPI) and/or Node.js/TypeScript (Next.js, Vercel AI SDK). Experience with Golang is a bonus.
• Practical experience in building evaluations, tracking quality metrics, and troubleshooting non-deterministic outputs in production.
• Familiarity with cost optimization strategies: model routing, caching, token usage monitoring, and prompt compression.
• Strong foundational knowledge in data structures, algorithms, and system design.
• Experience with containerized deployments (Docker, Kubernetes) and cloud platforms (AWS/GCP). A practical understanding of Kubernetes concepts and trade-offs is required.
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
Get handpicked remote jobs straight to your inbox weekly.