Applied AI Engineer

Posted 16 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Design and conduct randomized experiments related to tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom.

• Take ownership of the analytical component of the AI measurement program, which includes work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates.

• Develop and validate LLM-as-judge and classification pipelines utilizing sampling strategies, hand-labeled ground truth, precision and recall measurement, and revalidation.

• Expand AI capabilities beyond code authoring to encompass testing, environment and data setup, migration and modernization, code review, security remediation, and the assembly of certification evidence.

• Collaborate closely with constrained teams to pinpoint delivery bottlenecks and align AI capabilities accordingly.

• Create evaluations for internal AI capabilities, which include golden sets, regression suites, groundedness scoring, answer-quality scoring, cost telemetry, and latency telemetry.

• Identify effective practitioner behaviors, document and teach best practices, and prioritize the publication of practices over rankings.

• Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry.

• Communicate findings, limitations, delivery constraints, and cost/value implications to engineering leadership and finance.

• Take charge of the analytical and applied aspects of the measurement program while collaborating with a data engineer responsible for extraction, identity, and metric pipelines.


⛳️ Requirements

• 7+ years of experience in software engineering and quantitative analysis.

• Practical experience with LLM applications, including prompting, tool and function calling, context management, and evaluation.

• Proficiency in experimental design and causal inference, encompassing randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models.

• Strong skills in Python and SQL.

• Experience with libraries such as pandas, statsmodels, scikit-learn, or R.

• Proficient in instrumenting and extracting data from operational systems and APIs.

• Knowledge in sampling design that can withstand scrutiny.

• Experience with resolving identity across systems, integrating disparate data sources, and modeling summary tables.

• Capable of conducting exploratory analysis, distributions, cohort analysis, time-series analysis, and reporting on coverage and limitations.

• Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategy, and release and change management.

• Expertise in Git and GitLab instrumentation, including merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks.

• Experience with Jira and Confluence integration, including REST APIs, changelog and page version history, GitLab–Jira development panel, and field and label conventions.

• Careful handling of personnel-adjacent data and default aggregate reporting.

• Ability to communicate effectively with both executive and engineering audiences.

• Preferred: Experience with MCP servers and clients or similar connector frameworks.

• Preferred: Familiarity with agent frameworks such as LangGraph, LangChain, Bedrock Agents, or Strands.

• Preferred: Knowledge of enterprise deployment and telemetry of coding assistants.

• Preferred: Experience with server-side Git hooks, GitLab CI, and system or webhook-driven capture.

• Preferred: Experience with Confluence and Jira as MCP-connected systems.

• Preferred: Familiarity with evaluation tools such as Ragas, DeepEval, or Bedrock model evaluation.

• Preferred: Experience with LLM observability tools such as LangFuse, Arize, or OpenTelemetry-based tracing.

• Preferred: Knowledge of Amazon Bedrock and AWS cost and usage data.

• Preferred: Familiarity with engineering productivity frameworks such as DORA, DX Core 4, or SPACE.

• Preferred: Background in program analysis, test generation, or developer tooling research.

• Preferred: Experience with dbt, Airflow, Dagster, or equivalent transformation and orchestration tools.

• Preferred: Knowledge of warehouse or lakehouse modeling.

• Preferred: Experience with BI and visualization tools.

• Preferred: Knowledge of queueing and flow analysis.


🏝️ Benefits

• Permanent full-time employment.

• Remote work flexibility.

• Global and multicultural environment with collaboration opportunities across US, Portugal, India, and Singapore offices.

• A dynamic, fast-paced, and impact-driven work atmosphere.

• Ownership mindset: engineers take responsibility for what they build.

• A collaborative, friendly, open, curious, and supportive culture.

People also viewed

CodiLime3 hours ago

AI Engineer – Fullstack

PL flagPoland OnlyFreelanceAI EngineerPLN 16.5k – PLN 28k/month
ApplyView job
GFT Technologies4 hours ago

Senior ML/AI Engineer

BR flagBrazil OnlyFull-timeAI Engineer
ApplyView job
GFT Technologies8 hours ago

Mid-Level ML/AI Engineer

BR flagBrazil OnlyFull-timeAI Engineer
ApplyView job
CareSource9 hours ago

Data Solutions AI Application Developer III

US flagUnited States OnlyFull-timeAI Engineer$94.1k – $164.8k/year
ApplyView job
Empath Health9 hours ago

AI Developer

US flagFlorida OnlyFull-timeAI Engineer
ApplyView job
eSimplicity16 hours ago

Principal AI Engineer

US flagMaryland OnlyFull-timeAI Engineer$143.6k – $205k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers