
AI/ML Ops Engineer
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Canada.
• Take ownership of the AI/ML loop from start to finish at scale across all pipelines, encompassing model training, deployment, monitoring, and retirement.
• Create, enhance, and implement ML models.
• Design, construct, and manage the infrastructure for model building and serving.
• Deploy trained scripts as live endpoints capable of handling real-time requests.
• Execute ML workflows as containerized Infrastructure as Code using tools like Terraform, GitHub Actions, Docker, and Kubernetes.
• Develop and automate standardized container pipelines for training, feature engineering, and inference processes.
• Oversee CI/CD practices using GitHub.
• Establish a comprehensive testing strategy throughout the entire ML pipeline, which includes model validation, integration, load, and deployment testing.
• Create visibility and alerting mechanisms for deployed pipelines.
• Design ML governance tools for the oversight and management of deployed infrastructure.
• Apply best practices in data, feature, and model lifecycle management.
• Play a significant role in AI architecture and design decisions, with a primary focus on ML pipeline initiatives.
• Collaborate with Engineering, the Security Operations Center (SOC), and the Adversary Pursuit Group (APG).
• Report directly to the Vice President of AI and Data.
• A minimum of 5 years of practical ML Engineering experience.
• Experience in personally training and deploying models in a production setting.
• A well-architected approach with an emphasis on efficiency, performance, security, and reliability.
• Comfortable managing deployment pipelines from start to finish.
• Strong analytical and problem-solving skills.
• Proficient in data-driven decision-making.
• Excellent communication and interpersonal abilities.
• Capable of influencing and collaborating with stakeholders at all levels.
• Experience with cloud-based ML infrastructure, particularly AWS.
• Familiarity with SageMaker and Bedrock.
• Experience with Kafka and Spark for managing inference streams and event-driven processing.
• Proficient in MLflow and SageMaker Pipelines.
• Knowledge of Terraform, AI CI/CD, and GitHub Actions.
• Familiarity with ML governance, including data, model, and feature versioning, monitoring, and testing.
• Experience with Docker, Kubernetes, and ECS/EKS.
• Proficient in Python and Bash.
• Proficient in SQL and SparkSQL.
• Knowledge of GitFlow, CI/CD workflows, and DevOps best practices.
• Experience with AI-assisted development lifecycles.
• Proven track record of building high-availability, production-grade systems with visibility and alerting.
• Nice to have: Experience with Transformer Neural Networks.
• Nice to have: Familiarity with Agile Scrum/Kanban methodologies.
• Nice to have: Experience with Anthropic, OpenAI, and LiteLLM APIs and SDKs.
• Nice to have: Background in Cybersecurity, IoT, or NLP.
• Nice to have: Experience with Grafana or CloudWatch.
• Global equity participation available for employees, with program details varying based on location and employment structure.
• Eligibility for a discretionary bonus.
• Competitive benefits for international employees in alignment with local market standards and applicable country laws.
• Eligible US employees will receive Health Insurance, Vision, Dental, and Life Insurance plans.
• Eligible US employees benefit from a robust 401k plan.
• Eligible US employees are granted Discretionary Time Off.
• Additional minor perks included.
Invisible Technologies
Combine | Global Recruitment
Your Software Supplier
Spotify
Get handpicked remote jobs straight to your inbox weekly.