
Senior Machine Learning Engineer – MLOps
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in India.
• Create, develop, and maintain deployment pipelines that transition models from development through validation into production.
• Implement model versioning, lineage, registry, and automated promotion methodologies.
• Establish consistent production-readiness standards and deployment practices for ML and AI workloads.
• Collaborate with Data Scientists to ensure smooth, consistent, and production-ready model handoffs.
• Take responsibility for monitoring production concerning model performance, drift, data quality, inference health, latency, and availability.
• Set up alerts and operational thresholds to detect degradation before it affects products or customers.
• Troubleshoot production failures, conduct root-cause analysis, and execute durable corrective measures.
• Develop operational practices that enhance reliability as the production model portfolio expands.
• Deploy and provide support for production LLM applications, including retrieval-based and agentic frameworks.
• Create evaluation frameworks for assessing the quality, reliability, and performance of generative AI.
• Track token consumption, inference costs, and cost-per-interaction metrics.
• Enforce controls regarding model access, usage, safety, and production behavior.
• Design and manage infrastructure for model training, validation, and retraining.
• Build automated retraining pipelines activated by performance, data, or business conditions.
• Ensure that training environments and workflows are reproducible, scalable, and observable.
• Collaborate with Data Engineering and Data Science to guarantee reliable data movement throughout the ML lifecycle.
• Implement model access controls, auditability, lineage, and governance standards.
• Assist with model risk classification and appropriate controls based on use cases.
• Generate documentation and technical evidence for security, compliance, and internal governance.
• Address incidents, troubleshoot failures, and coordinate resolutions across teams.
• Develop runbooks and operational procedures that minimize reliance on tribal knowledge.
• Identify recurring operational challenges and automate solutions where feasible.
• Work independently while collaborating with U.S.-based Data Science and Data Engineering teams.
• 6+ years of experience in software engineering, data engineering, machine learning engineering, or a related technical field.
• 3+ years of hands-on experience in deploying and managing AI/ML systems in a production environment.
• Proven experience supporting both traditional machine learning and LLM-based workloads in production settings.
• Strong grasp of the entire model lifecycle, including development, validation, deployment, monitoring, retraining, and retirement.
• Production experience with LLM-powered applications and agentic frameworks.
• Familiarity with retrieval architectures, evaluation methodologies, and production monitoring for generative AI.
• Understanding of LLM performance, latency, token usage, and cost-per-interaction management.
• Ability to implement practical operational and governance controls around generative AI systems.
• Extensive experience with a major cloud platform and its managed AI/ML services; AWS is highly preferred.
• Practical experience with model registries, pipeline orchestration, ML CI/CD, automated retraining, and production monitoring.
• Strong Python programming skills.
• Experience with containerization and infrastructure-as-code methodologies.
• Ability to design reliable, repeatable, and automated production environments.
• Experience managing production services with significant responsibility for reliability and availability.
• Excellent incident response, troubleshooting, and root-cause analysis capabilities.
• Ability to differentiate symptoms from underlying system failures and implement long-term fixes.
• Comfortable making sound operational decisions independently when immediate U.S.-based support is unavailable.
• Strong written communication skills for technical documentation.
• Experience generating runbooks, architectural documentation, standards, and operational procedures.
• Proactive communication style appropriate for distributed, asynchronous teams.
• Ability to collaborate effectively across Data Engineering, Data Science, Product, and other technical teams.
• Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
• Experience in implementing AI governance, model risk tiering, or access-control frameworks is a plus.
• Familiarity with modern data warehousing and orchestration technologies in production analytics environments is advantageous.
• Experience supporting large-scale data and ML workloads within an AWS ecosystem is a plus.
• Successful experience working in distributed global teams alongside U.S.-based colleagues is a plus.
• Competitive local benefits offered through our Employer of Record.
• Flexible remote working environment.
• Continuous professional development opportunities.
• Chance to collaborate directly with U.S.-based Data and Technology teams.
• Significant ownership of production systems that support Dynatron’s AI strategy.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.