
Senior Software Engineer
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Texas.
β’ Oversee the design and advancement of Pearson's Kubernetes-driven machine learning platform that facilitates large-scale model training and deployment.
β’ Create, implement, and enhance distributed machine learning workflows utilizing MetaFlow and various cloud-native technologies.
β’ Develop platform functionalities for reproducible experimentation, automated model training, artifact management, and deployment in production.
β’ Build infrastructure that supports GPU-based workloads for conventional machine learning models, foundational models, and agentic pipelines.
β’ Design and execute backend services and APIs that facilitate machine learning lifecycle management.
β’ Assess and incorporate open-source technologies to enhance developer productivity, platform reliability, scalability, and operational efficiency.
β’ Work alongside AI scientists to transform research prototypes into resilient, scalable, production-grade systems.
β’ Enhance platform observability, reliability, security, and efficiency in cloud cost management.
β’ Guide engineers, contribute to the technical strategy, and establish engineering best practices throughout the team.
β’ A Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field, or equivalent professional experience.
β’ Significant software engineering experience in developing complex distributed systems.
β’ Proficient in Python development at an expert level.
β’ Experience in designing and implementing cloud-native applications on AWS.
β’ Familiarity with developing applications utilizing Kubernetes and container technologies.
β’ Experience in designing RESTful APIs and microservice architectures.
β’ Proficient in working with both SQL and NoSQL databases.
β’ Familiarity with CI/CD pipelines, Git-based development workflows, and automated testing.
β’ Strong problem-solving, communication, and collaboration abilities.
β’ Preferred: Experience with machine learning platforms such as MetaFlow, MLflow, Kubeflow, or comparable workflow orchestration systems.
β’ Preferred: Background in production machine learning systems.
β’ Preferred: Experience with GPU computing and distributed model training.
β’ Preferred: Knowledge of large language model deployment or inference infrastructure.
β’ Preferred: Familiarity with PyTorch, TensorFlow, or similar machine learning frameworks.
β’ Preferred: Experience with Kubernetes operations, scheduling, and workload optimization.
β’ Preferred: Proficiency in Go development.
β’ Preferred: Familiarity with Infrastructure as Code technologies.
β’ Preferred: Skills in performance optimization and cloud cost management.
β’ Preferred: Experience in building internal developer platforms or engineering productivity tools.
β’ Experience creating platforms utilized by machine learning engineers and data scientists.
β’ Experience in deploying and managing production AI or LLM infrastructure.
β’ Experience in fine-tuning, deploying, and overseeing foundation models and pipelines.
β’ Experience in designing highly scalable cloud-native systems that manage large datasets and compute-intensive tasks.
β’ A curiosity about emerging AI technologies and the ability to evaluate them in a pragmatic manner.
β’ A passion for developing tools that empower others to accelerate their work.
β’ No eligibility for bonuses.
β’ Information regarding benefits can be found on the linked benefits page.
Arista Networks
Coforma
Platform.sh
Arista Networks
Get handpicked remote jobs straight to your inbox weekly.