
Machine Learning Engineer β ML Training Platform
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in United States, +1 more country.
β’ Design, develop, and enhance the platform that facilitates ongoing experimentation and extensive training on consumer nodes and cloud instances.
β’ Create resource management systems for provisioning and orchestrating computing across AWS, GCP, and Azure utilizing Pulumi/Terraform.
β’ Manage dynamic scaling, state synchronization, and concurrent operations across numerous diverse nodes.
β’ Engineer a fault-tolerant distributed machine learning infrastructure for GPU clusters using NVIDIA runtime.
β’ Implement S3 checkpointing, manage large datasets and streaming, monitor health, and develop resilient retry strategies.
β’ Construct systems that simulate and manage bandwidth shaping, latency injection, and packet loss.
β’ Oversee node churn and ensure continuous data flow across workers with varying connectivity.
β’ Proven production experience with infrastructure-as-code tools (Pulumi/Terraform/CloudFormation) in managing multi-cloud deployments.
β’ Proficiency with Docker/Kubernetes (EKS), GPU workloads, and large-scale heterogeneous clusters.
β’ Knowledge of distributed training workflows, including checkpointing, data sharding, model versioning, and orchestration of long-running jobs.
β’ Familiarity with decentralized networking concepts, such as P2P, NAT traversal, traffic shaping, and real bandwidth limitations.
β’ Strong skills in Python engineering, particularly with asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling.
β’ Practical experience in observability and SRE practices, including Prometheus/Grafana, performance profiling, and incident management.
β’ Background in a startup environment with significant service orchestration experience or at a large tech company.
β’ Capability to showcase the systems you have owned.
β’ A belief in Protocol Learning as a promising approach for collective, trustless, and sovereign AI.
β’ Professional-level proficiency in English, both written and spoken.
β’ Willingness to collaborate across various time zones.
β’ Substantial equity/ownership opportunities for key technical contributors, in addition to a competitive base salary.
β’ Flexible work environment with a globally distributed team.
β’ Option for full visa sponsorship and relocation assistance to either Australia or the US.
β’ Chance to engage with innovative, unpublished challenges in the training and serving of frontier models.
Sourcegraph
Quora
NBCUniversal
Spotify
Get handpicked remote jobs straight to your inbox weekly.