
Technical Lead
Posted Aug 20

Posted Aug 20
This is a fully remote position, open to applicants in Poland.
• Act as the inaugural technical leader for Veris EvalOps
• Develop the platform from the initial line of code to the first paying clients within 6–7 months
• Design and implement a pre-production Release Gate that generates reproducible AI readiness scores
• Create and deploy a Knowledge Health Monitor that continuously evaluates enterprise AI knowledge bases
• Oversee the evaluation engine, which includes LLM-as-judge scoring, rule-based validations, groundedness verification, hallucination detection, and regression comparisons
• Establish distributed tracing and observability for LLM calls, RAG retrievals, and agent workflows utilizing OpenTelemetry
• Track token usage, tool invocations, costs, and latency metrics
• Develop the knowledge health pipeline for ingestion, detection of stale content, contradiction analysis, and coverage-gap mapping
• Manage multi-tenant platform architecture, RBAC, API connectors, dashboards, and CI/CD integrations
• Recruit and manage a team of 5–7 engineers across Platform Core, Release Gate, and Knowledge Health
• Make comprehensive architecture decisions
• Ensure the first pilot client is operational by month 3
• Collaborate directly with pilot clients during onboarding and review of results
• Over 5 years of experience in engineering
• At least 2 years of experience leading a team of 3–8 through a full build cycle from architecture to deployment to paying users
• Proficiency in LLM evaluation methodologies, including LLM-as-judge design, RAGAS/DeepEval-style metrics, construction of golden datasets, regression testing for AI systems, and hallucination detection
• Familiarity with OpenTelemetry for LLM observability and tracing across LLM calls, RAG retrievals, and multi-step agent processes
• Experience in cost/latency attribution, drift detection, and anomaly detection
• Hands-on experience with full-pipeline RAG system architecture, encompassing chunking, embeddings, vector stores, retrieval, and re-ranking
• Practical experience with AI agent patterns including ReAct, Plan-and-Execute, and supervisor/sub-agent architectures
• Knowledge of tool-call evaluation and guardrails
• Proficient in production-grade async Python, FastAPI, Celery, multi-tenant SaaS architecture, PostgreSQL/Redis, and CI/CD integration
• Capability to assess build-versus-integrate decisions, including Langfuse/Braintrust versus developing from scratch
• Comfort in client-facing roles and ability to engage directly with pilot clients
• Experience with multi-provider LLMs such as OpenAI, Anthropic, Azure OpenAI, or Bedrock (preferred)
• Background in MLOps and experiment tracking using MLflow, Weights & Biases, or similar tools (preferred)
• Understanding of the EU AI Act and NIST AI RMF, PII management in AI pipelines, and basics of red-teaming (preferred)
• Flexibility for remote work
• Chance to develop a zero-to-one AI evaluation platform
• Opportunity to hire and lead a team of 5–7 engineers
• Direct influence over architectural decisions
• Engage with pilot clients during onboarding and results evaluation
Webflow
RecruityTalent
Viceri SEIDOR
Ponta
Get handpicked remote jobs straight to your inbox weekly.