
Principal Machine Learning Engineer
Posted 13 hours ago

Posted 13 hours ago
This is a fully remote position, open to applicants in United States, +1 more state.
• Take full responsibility for the machine learning platform, overseeing everything from data and feature pipelines to training infrastructure, model registry and lineage, inference services, and deployment.
• Create and develop integrations with the broader Accelerant platform, third-party providers, and systems managed by other teams.
• Streamline deployment processes through versioning, staged rollouts, rollbacks, and continuous integration/continuous deployment (CI/CD) for models and agents.
• Develop monitoring systems that identify data drift, pipeline failures, and actual performance degradation, even in cases where labels are delayed.
• Construct the infrastructure necessary for agentic AI, which includes orchestration, tool and API integration, retrieval, caching, and management of cost and latency controls.
• Ensure reliability, cost-effectiveness, and performance across machine learning workloads, from batch scoring to low-latency services.
• Establish model governance and maintain audit trails for regulators and internal risk committees.
• Lead and expand the function by setting technical standards, mentoring a small team, and collaborating with data scientists.
• Extensive experience in managing machine learning systems in production, including post-launch operations.
• Strong engineering background in Python, infrastructure as code, containers, and orchestration.
• Proficiency in at least one major cloud service provider.
• Good judgment regarding cost implications and potential failure scenarios.
• Data engineering skills across pipelines, orchestration, storage, and access patterns.
• Adequate SQL skills to function effectively within a data warehouse environment.
• Experience in integrating systems across organizational boundaries and influencing teams without direct authority.
• Statistical knowledge sufficient for evaluating model performance alongside data scientists.
• Experience leading or mentoring engineers.
• Sound judgment regarding infrastructure investments and operational value.
• Willingness to engage with large language models (LLMs) and agentic AI.
• Strong communication skills and credibility to assess when a system is not prepared for deployment.
• Experience with LLM/agentic infrastructure, regulated industries, insurance or financial services, predictive modeling, internal platforms, real-time systems, streaming, feature stores, or high-throughput scoring is advantageous.
• Ownership of a function with the autonomy to determine its operation.
• A team of capable data scientists who will benefit from your contributions.
• Opportunities to tackle challenges ranging from overnight batch scoring to low-latency services and agentic systems.
• A collaborative group of individuals who enjoy working together to solve complex problems.
• High degree of autonomy with minimal bureaucracy.
Doma
CSC Generation
Capgemini
Airbnb
Get handpicked remote jobs straight to your inbox weekly.