Machine Learning Engineer – ML Training Platform

Posted Sep 1

This is a fully remote position, open to applicants in United States, +1 more country.

πŸ“‹ Description

β€’ Design, develop, and enhance the platform that facilitates ongoing experimentation and extensive training on consumer nodes and cloud instances.

β€’ Create resource management systems for provisioning and orchestrating computing across AWS, GCP, and Azure utilizing Pulumi/Terraform.

β€’ Manage dynamic scaling, state synchronization, and concurrent operations across numerous diverse nodes.

β€’ Engineer a fault-tolerant distributed machine learning infrastructure for GPU clusters using NVIDIA runtime.

β€’ Implement S3 checkpointing, manage large datasets and streaming, monitor health, and develop resilient retry strategies.

β€’ Construct systems that simulate and manage bandwidth shaping, latency injection, and packet loss.

β€’ Oversee node churn and ensure continuous data flow across workers with varying connectivity.


⛳️ Requirements

β€’ Proven production experience with infrastructure-as-code tools (Pulumi/Terraform/CloudFormation) in managing multi-cloud deployments.

β€’ Proficiency with Docker/Kubernetes (EKS), GPU workloads, and large-scale heterogeneous clusters.

β€’ Knowledge of distributed training workflows, including checkpointing, data sharding, model versioning, and orchestration of long-running jobs.

β€’ Familiarity with decentralized networking concepts, such as P2P, NAT traversal, traffic shaping, and real bandwidth limitations.

β€’ Strong skills in Python engineering, particularly with asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling.

β€’ Practical experience in observability and SRE practices, including Prometheus/Grafana, performance profiling, and incident management.

β€’ Background in a startup environment with significant service orchestration experience or at a large tech company.

β€’ Capability to showcase the systems you have owned.

β€’ A belief in Protocol Learning as a promising approach for collective, trustless, and sovereign AI.

β€’ Professional-level proficiency in English, both written and spoken.

β€’ Willingness to collaborate across various time zones.


🏝️ Benefits

β€’ Substantial equity/ownership opportunities for key technical contributors, in addition to a competitive base salary.

β€’ Flexible work environment with a globally distributed team.

β€’ Option for full visa sponsorship and relocation assistance to either Australia or the US.

β€’ Chance to engage with innovative, unpublished challenges in the training and serving of frontier models.

People also viewed

Sourcegraph21 hours ago

ML Engineer, Agentic Systems

North AmericaFull-timeMachine Learning Engineer$88k – $176k/year
ApplyView job
Quora23 hours ago

Senior Machine Learning Engineer, Ads

US flagUnited States OnlyFull-timeMachine Learning Engineer$189.5k – $274.6k/year
ApplyView job
NBCUniversal1 day ago

Staff MLOps Engineer

CA flagCanada OnlyFull-timeMachine Learning Engineer
ApplyView job
Spotify1 day ago

Staff Machine Learning Engineer – Home Surfaces

US flagNew York OnlyFull-timeMachine Learning Engineer
ApplyView job
Fullscript1 day ago

Senior Machine Learning Engineer

CA flagCanada OnlyFull-timeMachine Learning EngineerC$140k – C$160k/year
ApplyView job
Vida Health1 day ago

Principal AI/ML Engineering Lead

US flagUnited States OnlyFull-timeMachine Learning Engineer$250k – $275k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers