
Senior Software Engineer – Applied AI
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in United States.
• Design and conduct randomized experiments focusing on tier routing, MCP coverage, permission configurations, repository context quality, and budget headroom.
• Take charge of the analytical aspect of the measurement program, which includes work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates.
• Develop and validate LLM-as-judge and classification pipelines utilizing sampling, hand-labeled ground truth, precision and recall measurement, and revalidation processes.
• Enhance AI capabilities across testing, data setup, migration, modernization, code review, security remediation, and certification evidence assembly.
• Collaborate directly with constrained teams to pinpoint delivery bottlenecks and tailor capabilities accordingly.
• Create evaluations for internal AI capabilities, such as golden sets, regression suites, groundedness, answer-quality scoring, cost, and latency telemetry.
• Identify effective practitioner behaviors, document and disseminate practices, and publish methodologies rather than mere rankings.
• Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry.
• Present findings to engineering leadership and finance, including successes, delivery constraints, and the limitations of each claim.
• Over 7 years of experience in software engineering and quantitative analysis.
• Practical experience with LLM applications, including prompting, tool and function calling, context management, evaluation, and insights into model failures in practice.
• Expertise in experimental design and causal inference, covering randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models.
• Proficient in Python and SQL, along with a statistical stack such as pandas, statsmodels, scikit-learn, or R.
• Experience in data collection from operational systems and APIs, with a focus on robust sampling design.
• Skills in identity resolution across systems and complex data joins.
• Capable of conducting exploratory analysis, examining distributions, cohort and time-series analysis, and reporting on coverage and limitations.
• Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategies, release, and change management.
• Proficient with Git and GitLab, including instrumentation depth, merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks.
• Experience with Jira and Confluence integration, including REST APIs, changelogs, page version history, GitLab–Jira development panel, fields, and labels.
• Attention to personnel-adjacent data and aggregate reporting.
• Strong communication skills to engage with executive and engineering audiences.
• Prefer experience with MCP servers and clients or similar connector frameworks.
• Knowledge of agent frameworks such as LangGraph, LangChain, Bedrock Agents, Strands, or equivalents.
• Experience in enterprise deployment of coding assistants and telemetry.
• Familiarity with server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance.
• Experience with Confluence and Jira as MCP-connected systems, including permission propagation, scoped credentials, and audit logging.
• Proficiency in evaluation tools such as Ragas, DeepEval, or Bedrock model evaluation; LLM observability tools like LangFuse, Arize, or OpenTelemetry-based tracing.
• Knowledge of Amazon Bedrock and AWS cost and usage data.
• Familiarity with engineering productivity frameworks such as DORA, DX Core 4, or SPACE.
• Experience in program analysis, test generation, or developer tooling research.
• Proficient in dbt, Airflow, Dagster, or similar tools for transformation and orchestration; warehouse or lakehouse modeling.
• Skilled in BI and visualization tools, with a focus on summary-table-based reporting.
• Understanding of queuing and flow analysis, including utilization, batch economics, and constraint identification.
• Global & Multicultural – Diverse perspectives and collaboration across offices in the US, Portugal, India, and Singapore.
• Startup Energy – A fast-paced, impact-driven environment.
• Ownership Mindset – Engineers take responsibility for what they build.
• Collaborative & Friendly – An open, curious, and supportive culture.
Shield AI
Netflix
Travoom
HeroSoftware GmbH - Shopify Apps
Get handpicked remote jobs straight to your inbox weekly.