
Applied AI Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
• Design and conduct randomized experiments related to tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom.
• Take ownership of the analytical component of the AI measurement program, which includes work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates.
• Develop and validate LLM-as-judge and classification pipelines utilizing sampling strategies, hand-labeled ground truth, precision and recall measurement, and revalidation.
• Expand AI capabilities beyond code authoring to encompass testing, environment and data setup, migration and modernization, code review, security remediation, and the assembly of certification evidence.
• Collaborate closely with constrained teams to pinpoint delivery bottlenecks and align AI capabilities accordingly.
• Create evaluations for internal AI capabilities, which include golden sets, regression suites, groundedness scoring, answer-quality scoring, cost telemetry, and latency telemetry.
• Identify effective practitioner behaviors, document and teach best practices, and prioritize the publication of practices over rankings.
• Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry.
• Communicate findings, limitations, delivery constraints, and cost/value implications to engineering leadership and finance.
• Take charge of the analytical and applied aspects of the measurement program while collaborating with a data engineer responsible for extraction, identity, and metric pipelines.
• 7+ years of experience in software engineering and quantitative analysis.
• Practical experience with LLM applications, including prompting, tool and function calling, context management, and evaluation.
• Proficiency in experimental design and causal inference, encompassing randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models.
• Strong skills in Python and SQL.
• Experience with libraries such as pandas, statsmodels, scikit-learn, or R.
• Proficient in instrumenting and extracting data from operational systems and APIs.
• Knowledge in sampling design that can withstand scrutiny.
• Experience with resolving identity across systems, integrating disparate data sources, and modeling summary tables.
• Capable of conducting exploratory analysis, distributions, cohort analysis, time-series analysis, and reporting on coverage and limitations.
• Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategy, and release and change management.
• Expertise in Git and GitLab instrumentation, including merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks.
• Experience with Jira and Confluence integration, including REST APIs, changelog and page version history, GitLab–Jira development panel, and field and label conventions.
• Careful handling of personnel-adjacent data and default aggregate reporting.
• Ability to communicate effectively with both executive and engineering audiences.
• Preferred: Experience with MCP servers and clients or similar connector frameworks.
• Preferred: Familiarity with agent frameworks such as LangGraph, LangChain, Bedrock Agents, or Strands.
• Preferred: Knowledge of enterprise deployment and telemetry of coding assistants.
• Preferred: Experience with server-side Git hooks, GitLab CI, and system or webhook-driven capture.
• Preferred: Experience with Confluence and Jira as MCP-connected systems.
• Preferred: Familiarity with evaluation tools such as Ragas, DeepEval, or Bedrock model evaluation.
• Preferred: Experience with LLM observability tools such as LangFuse, Arize, or OpenTelemetry-based tracing.
• Preferred: Knowledge of Amazon Bedrock and AWS cost and usage data.
• Preferred: Familiarity with engineering productivity frameworks such as DORA, DX Core 4, or SPACE.
• Preferred: Background in program analysis, test generation, or developer tooling research.
• Preferred: Experience with dbt, Airflow, Dagster, or equivalent transformation and orchestration tools.
• Preferred: Knowledge of warehouse or lakehouse modeling.
• Preferred: Experience with BI and visualization tools.
• Preferred: Knowledge of queueing and flow analysis.
• Permanent full-time employment.
• Remote work flexibility.
• Global and multicultural environment with collaboration opportunities across US, Portugal, India, and Singapore offices.
• A dynamic, fast-paced, and impact-driven work atmosphere.
• Ownership mindset: engineers take responsibility for what they build.
• A collaborative, friendly, open, curious, and supportive culture.
CodiLime
GFT Technologies
GFT Technologies
CareSource
Get handpicked remote jobs straight to your inbox weekly.