
Senior Machine Learning Engineer
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in United States.
• Assume operational responsibility for the machine learning and AI services deployed by Candid.
• Oversee service performance, manage retraining schedules, coordinate transitions from data scientists, and act as the primary contact for production models.
• Enhance inference efficiency for deployed models, including intricate graph inference models, through techniques such as quantization, artifact reduction, batching, and effective serialization.
• Design and manage systems for experiment tracking, model versioning, and artifact management.
• Develop and sustain observability for ML and AI services via centralized logging, metrics, dashboards, and alerts.
• Create and refine a repeatable AWS deployment process for new ML services, incorporating CI/CD integration and infrastructure-as-code methodologies.
• Track and manage AWS expenditures across ML workloads, including Bedrock token utilization, compute sizing, and S3 lifecycle management.
• Collaborate with data scientists to comprehend model behavior, extract operational insights, and convert research code into production-ready deployments.
• Act as a technical intermediary between Data Science and product/software engineering teams for ML and AI integration.
• Establish integration agreements, APIs, latency and reliability expectations, as well as input/output schemas.
• Develop and manage Amazon Bedrock-supported services and integrations.
• Contribute to secure-by-default ML services through IAM configurations, secrets management, and compliance-oriented tagging.
• Engage in technical planning and roadmap discussions.
• A minimum of 4 years of professional software engineering experience.
• At least 2 years of experience in MLOps, ML engineering, or data science with a focus on production ML systems as a core responsibility.
• Strong expertise in Python, including the development of production-quality service code.
• Practical experience with experiment tracking and model lifecycle tools such as MLflow or Weights & Biases.
• Hands-on experience in deploying PyTorch models in a production environment.
• Familiarity with quantization, batching, ONNX Runtime, model serialization, container/artifact optimization, or cold-start mitigation on Lambda/Fargate.
• Experience in deploying and monitoring ML models in production, including managing model degradation, drift signals, retraining triggers, and artifact oversight.
• Working knowledge of AWS ML deployment services, including Lambda, ECS/Fargate, S3, IAM, and CloudWatch.
• Proven experience in building or managing CI/CD pipelines for ML services.
• A history of enhancing production reliability through observability and disciplined deployment practices.
• Ability to collaborate closely with data scientists and articulate operational decisions effectively.
• Capability to work across different teams in software or product engineering.
• Comfortable taking ownership of work independently.
• Excellent written and verbal communication skills.
• Willingness to undertake additional duties and special projects as required.
• Awareness and respect for racial, gender, sexual orientation, and cultural diversity.
• Commitment to Candid's core values: driven, direct, accessible, curious, and inclusive.
• Health insurance (medical, dental, vision).
• Retirement savings plan with an additional matching option.
• Paid life insurance and accidental death & dismemberment coverage.
• Paid time off (PTO, compassionate leave, volunteer time, holidays, parental leave).
• Short-term and long-term disability insurance.
• Pre-tax transit benefits.
• Flexible spending accounts.
• Supplemental insurance options.
• Summer hours.
• Eligible employer for the Public Service Loan Forgiveness (PSLF) program.
• Remote work opportunities.
• Annual weeklong all-staff summits.
Sourcegraph
Quora
NBCUniversal
Spotify
Get handpicked remote jobs straight to your inbox weekly.