
Staff ML Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
β’ Design, construct, and launch production-ready AI agents and agentic workflows that are integrated with enterprise applications and business processes.
β’ Develop evaluation, monitoring, guardrail, and safety frameworks for applications powered by large language models (LLMs).
β’ Architect and implement scalable AI-driven capabilities within the current Java/Spring platform and distributed microservices environment.
β’ Establish engineering best practices for AI Agent systems, covering observability, testing, deployment, versioning, and operational excellence.
β’ Lead architectural decisions across AI services, microservices, APIs, and integrations with enterprise applications.
β’ Collaborate with AI Scientists and ML Engineers to bring models into production and convert experimental work into reliable software.
β’ Mentor engineers, conduct architectural reviews, and shape technical strategy across various teams.
β’ Assess emerging AI technologies and propose pragmatic adoption strategies to enhance product capabilities and developer productivity.
β’ Bachelor's degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
β’ Over 8 years of professional software engineering experience, including leading intricate technical initiatives in distributed systems.
β’ Strong expertise in Python, Java, and the Spring ecosystem, including Spring Boot and Spring Framework.
β’ Experience in building and evolving enterprise-scale applications.
β’ Profound knowledge in creating scalable, production-grade AI-enabled applications and distributed systems.
β’ Experience in architecting and deploying AI agents, agentic workflows, and LLM-powered applications at an enterprise level.
β’ Practical experience in integrating AI capabilities into existing enterprise applications.
β’ Extensive experience in designing and deploying production-quality LLM applications using technologies such as LangChain, LangGraph, Spring AI, Model Context Protocol (MCP), or similar.
β’ Strong understanding of Retrieval-Augmented Generation (RAG), tool calling, MCP, prompt engineering, and agent orchestration.
β’ Expertise in observability, CI/CD, automated testing, and production operations for AI applications.
β’ Proven ability to navigate ambiguity, mitigate technical risk, and influence architectural direction across cross-functional teams.
β’ Excellent communication, collaboration, and mentoring abilities.
β’ Preferred experience with Databricks, vector search technologies, embedding models, conversational AI, and evaluation strategies for AI agents.
β’ Preferred experience in building AI infrastructure, developer tooling, and platform capabilities.
β’ Preferred experience in architecting and maintaining large-scale Java microservices using Spring Boot.
β’ Preferred experience with AWS, Azure, or GCP, Kubernetes, and containerized deployments.
β’ Preferred familiarity with model serving, inference optimization, caching strategies, and cost optimization for enterprise AI systems.
β’ Competitive benefits package.
β’ Discretionary bonus or commission linked to achieved results.
β’ Reasonable accommodations for qualified individuals with disabilities or disabled veterans during the hiring process.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.