
Senior Platform Engineer, ML Infrastructure
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and manage scalable machine learning infrastructure and platform capabilities that support the comprehensive machine learning lifecycle across both cloud and on-premises settings.
• Create developer tools, services, and infrastructure that empower ML and engineering teams to efficiently build, deploy, and manage production systems.
• Take the lead on complex technical projects independently, from defining problems and designing architecture to implementing solutions, rolling out to production, and maintaining operational oversight.
• Make architectural and engineering choices that balance immediate delivery needs with long-term goals for scalability, reliability, and ease of maintenance.
• Develop reliable, scalable, and user-friendly platform capabilities that enhance developer productivity and streamline operations.
• Collaborate with ML engineers, infrastructure engineers, and stakeholders to convert customer requirements into effective platform solutions.
• Address infrastructure and platform challenges related to performance, reliability, scalability, and developer experience.
• Promote adoption and continuous enhancement through feedback from engineering teams.
• Set high benchmarks for software quality, operational excellence, and readiness for production.
• Provide platform solutions that deliver measurable engineering and business impacts across various teams and use cases.
• Over 5 years of professional experience in software engineering, emphasizing platform, infrastructure, or distributed systems.
• Proficient in Python engineering, particularly in developing production services, SDKs, automation, or platform tools.
• Proven experience in designing, building, and managing production platform capabilities utilized by multiple engineering teams.
• Familiarity with ML platform architecture and the complete ML lifecycle, including experimentation, distributed training, model deployment, and operational processes.
• Experience in building and managing applications on Kubernetes and cloud platforms (preferably AWS), with knowledge of production reliability, observability, and operational best practices.
• Strong technical judgment and the capacity to independently lead intricate technical projects from inception to production, collaborating effectively with ML engineers, infrastructure teams, and product stakeholders.
• Visa sponsorship is available for this role.
• Annual performance bonus.
• Competitive benefits package.
• Programs for recruiting, mentorship, career growth, and learning & development.
• Reasonable accommodations for individuals with disabilities.
• An inclusive workplace that values and promotes diversity.
Parallel Partners
Playbypoint
Alectrona
Cloudera
Get handpicked remote jobs straight to your inbox weekly.