
Applied AI Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Portugal.
• Develop and manage a production Slack-native SQL BI analyst agent that converts natural-language business inquiries into governed SQL.
• Validate queries, ensure result accuracy, and provide referenced evidence for every statistic.
• Create executive P&L responses, daily health reports, and data-driven root cause analyses.
• Enhance agents for customer-facing intent triage, routing, RAG responses, conversational context, localized brand voice, and escalation processes.
• Integrate agents with helpdesk and CRM systems via webhooks, session lifecycle management, intent tagging, and automated escalation tickets.
• Develop risk-stratified tools for back-office APIs that include validation, confirmation workflows, and controlled autonomy.
• Safeguard agents against prompt injection, tool misuse, and data exfiltration.
• Construct agents utilizing LangGraph, Anthropic Agent SDK, MCP, or comparable frameworks.
• Engineer feedback loops, semantic memory, decision audit logging, evaluation harnesses, and regression suites.
• Create and implement models for churn, lifetime value, bonus sensitivity, player risk, collusion, bot play, multi-accounting, and treasury/payment anomalies.
• Deploy governed ML signals with versioning, SLAs, freshness, drift, calibration monitoring, and automated retraining.
• Manage multi-vector withdrawal risk scoring and evidence-aware re-scoring.
• Convert policies into deterministic, configurable, auditable rules; simulate and backtest modifications.
• Design holdouts and control groups, assess uplift, and conduct in-depth analyses.
• Develop agent supervisor and approval interfaces, review queues, session replay, grading modules, and model/evaluation datasets.
• Create dashboards for decision audits, agent performance, risk review queues, and KPIs.
• Report directly to the Head of Data and collaborate with AI, BI, Infrastructure, Customer Success, and product teams.
• A minimum of 4 years' experience in software, data science, or machine learning engineering.
• At least 1 year of experience building LLM-powered agents in a production environment.
• Familiarity with tool usage and function calling, structured outputs, retrieval and memory, and multi-step orchestration.
• Experience with LangGraph, Anthropic Agent SDK, MCP, or similar orchestration frameworks.
• Background in production RAG systems, including grounding, chunking, retrieval quality, hallucination control, and refusal strategies.
• Understanding of prompt injection, tool-call abuse, and data leakage defenses.
• Ownership of the production ML lifecycle, including feature engineering, training, serving, monitoring, and retraining.
• Experience in fraud, risk, or abuse detection is highly desirable.
• Statistical rigor in experimental design, including holdouts, control groups, uplift measurement, and score calibration.
• Experience in constructing evaluation harnesses and regression suites for non-deterministic systems.
• Proficient in Python.
• Comfortable working with TypeScript.
• Strong SQL skills, particularly with columnar analytical databases; ClickHouse is preferred.
• Capability to build stakeholder-ready interfaces using a front-end framework or tools like Streamlit.
• Experience designing systems where model outputs facilitate deterministic execution.
• Familiarity with LLM observability and tracing tools, such as Langfuse or LangSmith.
• Nice to have: experience in iGaming or high-trust transaction-intensive environments.
• Nice to have: knowledge of helpdesk or customer service platform integration.
• Nice to have: experience with blockchain or crypto-native transaction flows.
• Nice to have: expertise in constrained optimization, bandits, or reinforcement learning.
• Nice to have: knowledge of rule engines, decision management systems, and Slack app development.
• Nice to have: experience with Kafka or MSK consumers, idempotent processing, and failure handling.
• Fully remote working environment.
• Asynchronous-first communication culture.
• Preference for EU timezone overlap.
• High level of autonomy and ownership over domain decisions.
• Architecture decisions are documented and actively debated.
• Opportunity to work in a global, multi-tenant scale engineering environment.
Alzheimer's Association®
Capital One
Capital One
Get handpicked remote jobs straight to your inbox weekly.