
Senior AI Engineer, Agents
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Greece, +4 more countries.
• Develop and sustain the agent roadmap by pinpointing repetitive expert tasks in collaboration with sales, data, and product teams.
• Evaluate potential agents based on feasibility, cost, and impact.
• Architect agent designs and create production code for tools, loops, sub-agents, memory, state management, and evaluations.
• Construct a reusable platform layer over EveryWatch data, encompassing pricing, auctions, listings, references, and portfolios.
• Create an MCP interface atop the current backend.
• Establish shared memory, tracing mechanisms, and reusable evaluation infrastructure.
• Enhance WatchChat's multi-turn state, memory management, latency, cost-efficiency, tool-call reliability, and regression checks.
• Generate high-quality multi-turn datasets, programmatic tool/argument verifiers, and LLM-as-judge evaluations on held-out datasets.
• Streamline model selection, context budgets, caching strategies, and fine-tuning to effectively manage costs and latency.
• Collaborate with engineering, product, sales, and data teams to identify and implement AI capabilities for collectors, dealers, sales teams, data operations, and engineering.
• Over 6 years of experience delivering production software, including at least 2 years working with LLM systems utilized by actual users.
• Proficient in ReAct or similar loops, tool/function invocation, planning, state and checkpointing, long-term memory, HITL processes, and streaming.
• Practical experience with LangGraph or a compelling case for an alternative solution.
• Familiarity with tracing and observability stacks such as LangSmith or Langfuse.
• Strong background in production Python, including FastAPI, asynchronous job management, and clear service boundaries.
• Solid experience with proper RAG: retrieval, re-ranking, relevance assessment, and evaluation.
• Knowledge of RAGAS, DeepEval, GEVAL, LLM-as-judge, or pass@k evaluation methodologies.
• Cloud production experience with AWS (Bedrock, SQS, EC2/EKS, S3) or a comparable platform.
• Experience with Docker and CI/CD processes.
• Strong architectural insight and the ability to question existing technical choices.
• Familiarity with LLM post-training methods such as SFT, LoRA/QLoRA, preference/RL techniques, reward design, or verifier design is a plus.
• Experience in voice agents, vision-language projects, multi-agent coordination, MCP integrations, text-to-SQL, report/document generation agents, guardrails, PII management, hallucination detection, or risk scoring is advantageous.
• A passion for watches, collectibles, or market data is a bonus, but not mandatory.
• Full ownership of EverWatch's entire agent layer and the roadmap associated with it.
• Access to a unique dataset that is unavailable elsewhere.
• Direct communication with the CTO; decisions made in days rather than quarters.
• Budget allocated for models, tools, and the coding agents you wish to utilize.
• Competitive salary package.
• Remote-first work arrangement.
GLOBALTALENT
GLOBALTALENT
knowmad mood
Tieto
Get handpicked remote jobs straight to your inbox weekly.