
Senior Principal Software Engineer, Machine Learning
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Take charge of the technical strategy for the ML Platform, encompassing feature store, model hosting and serving, experimentation, and training infrastructure.
• Influence architectural choices regarding scalability, reliability, latency, and cost-effectiveness.
• Spearhead the design and execution of extensive platform projects from initial concept to full production.
• Collaborate across ML, data, and infrastructure teams.
• Identify and address systemic issues related to online/offline feature parity, model deployment, experimentation speed, GPU utilization, and inter-team dependencies.
• Uphold a high standard of engineering quality through active code contributions, design evaluations, and mentorship.
• Work alongside ML engineering, data science, product, and platform leadership to align ML strategy with technical roadmaps.
• Establish secure pathways for model deployment, including feature registration, canary rollout, monitoring, and rollback processes.
• Utilize AI-enhanced development tools to boost development speed and enhance code quality.
• Over 10 years of experience in delivering complex backend or infrastructure systems at scale.
• Direct involvement in building or managing core ML infrastructure, such as feature stores, model serving, experimentation platforms, or training orchestration.
• Proficiency in a contemporary backend programming language, preferably Java or Kotlin.
• Extensive knowledge of distributed systems principles, including consistency, latency, throughput, fault tolerance, and observability.
• Strong grasp of data modeling, query languages, and the online/offline data patterns that support ML systems.
• Proven track record of technical leadership with the ability to foster cross-team alignment and influence engineering, product, and business stakeholders.
• Bachelor's degree in Computer Science or a related discipline, or equivalent hands-on experience.
• Desirable experience with ML platform components such as Tecton, MLflow, SageMaker, or Databricks.
• Desirable experience in creating experimentation or A/B testing platforms at scale.
• Desirable familiarity with Kafka, Flink, or Spark Streaming.
• Desirable experience in deploying LLMs or deep-learning models in production, GPU capacity planning, and inference optimization.
• Desirable experience in supporting internal developer-facing platforms with a product-oriented mindset.
• Competitive salary and benefits packages.
• Cash compensation that may include overtime and bonuses/commissions if applicable.
• Equity options available.
• Flexible work arrangements and a hybrid work model that respects individual needs.
• Access to AI tools across various disciplines to facilitate quicker, more independent, and higher-quality work.
• Reasonable accommodations provided for individuals with disabilities.
SumerSports
Airbnb
The Home Depot
Get handpicked remote jobs straight to your inbox weekly.