
Machine Learning Platform Engineer, Machine Learning, Artificial Intelligence
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in Mali, +1 more country.
β’ Develop and maintain the Machine Learning (ML) infrastructure and platforms that drive Artificial Intelligence (AI) products.
β’ Create systems for model training, evaluation, deployment, inference, and experimentation.
β’ Construct and enhance model serving and inference infrastructure to accommodate high-throughput and low-latency demands.
β’ Enhance the reliability, scalability, latency, and cost-effectiveness of Artificial Intelligence (AI) systems.
β’ Design dependable pipelines for data preparation, training, evaluation, model deployment, and ongoing enhancement.
β’ Create platforms and tools that empower Artificial Intelligence (AI) engineers and researchers to experiment, assess, and deploy models more efficiently.
β’ Develop infrastructure for evaluation and benchmarking to assess model quality, performance, and potential regressions.
β’ Implement production-level observability, monitoring, tracing, and alerting for Artificial Intelligence (AI)/Machine Learning (ML) applications.
β’ Identify performance bottlenecks within the Machine Learning (ML) stack and consistently enhance system efficiency.
β’ Collaborate closely with Artificial Intelligence (AI) engineers, researchers, and product teams to translate evolving model requirements into production-ready infrastructure.
β’ Ensure that AI infrastructure consistently supports production workloads at scale.
β’ Guarantee that models can be trained, assessed, deployed, and improved in an efficient manner.
β’ Ensure that inference systems provide excellent latency, throughput, reliability, and cost efficiency.
β’ Maintain that Machine Learning (ML) pipelines are reproducible, observable, maintainable, and robust.
β’ Swiftly detect and diagnose regressions in models and infrastructure.
β’ Create reusable primitives for the Machine Learning (ML) infrastructure platform.
β’ Facilitate the rapid evolution of the AI stack as new models, architectures, and inference techniques are developed.
β’ Proven experience in Machine Learning (ML) and Artificial Intelligence (AI).
β’ Solid foundation in software engineering principles and experience in building production systems.
β’ Experience in constructing Machine Learning (ML) infrastructure, platforms, or operational machine learning systems.
β’ Familiarity with model deployment, inference, evaluation, or data pipelines.
β’ Strong grasp of distributed systems and system reliability.
β’ Capability to write clean, maintainable, production-quality code.
β’ Comfortable navigating ambiguous and fast-paced environments.
β’ A proactive approach towards ownership, experimentation, and continuous improvement.
β’ Proficiency in Python, PyTorch, JAX, LLM, and ML serving infrastructures such as vLLM, SGLang, or TensorRT-LLM, along with cloud infrastructure, distributed systems, ML/data pipelines, workflow orchestration, GPU infrastructure and performance tools, and vector databases and retrieval systems.
β’ Willingness to complete a 60-minute coding assessment.
β’ Medical insurance.
β’ Dental insurance.
β’ Vision insurance.
β’ Savings Plan Options.
β’ Paid Time Off (PTO).
Playbypoint
Alectrona
Cloudera
DKSH Portugal, Unipessoal, Lda.
Get handpicked remote jobs straight to your inbox weekly.