Senior Machine Learning Engineer

Posted 19 hours ago

This is a fully remote position, open to applicants in Pennsylvania, +1 more state.

📋 Description

• Develop infrastructure that supports the development, evaluation, deployment, and ongoing enhancement of language models and AI systems.

• Take ownership of LLMOps, fine-tuning infrastructure, model evaluation, dataset pipelines, experiment management, model serving, and production observability.

• Create platforms that allow AI engineers and researchers to conduct experiments swiftly while ensuring reproducibility, scalability, and reliability.

• Assist in the complete model lifecycle from dataset creation and experimentation to training, evaluation, deployment, monitoring, and iteration.

• Design and construct LLMOps infrastructure tailored for production language models.

• Establish scalable training and fine-tuning infrastructure for both commercial and open-weight language models.

• Develop pipelines for supervised fine-tuning, parameter-efficient fine-tuning, preference optimization, and various post-training methodologies.

• Construct distributed training and GPU-accelerated machine learning infrastructure.

• Create data pipelines for training, fine-tuning, evaluation, and synthetic data generation.

• Implement dataset versioning, lineage, quality validation, transformation, and reproducible experimentation systems.

• Develop infrastructure for experiment management to compare models, datasets, hyperparameters, prompts, and training methods.

• Build automated model evaluation pipelines and checks for production readiness.

• Design model registries, artifact management, versioning, and promotion workflows.

• Construct and manage scalable model-serving and inference infrastructure.

• Create abstractions that support multiple models and inference providers.

• Implement observability for training and inference, covering metrics, tracing, logging, resource utilization, model quality, latency, throughput, and cost.

• Optimize workloads for GPU efficiency, throughput, latency, reliability, and infrastructure expenses.

• Establish automated workflows for deployment, rollback, canarying, and production validation.

• Investigate failures across data pipelines, training jobs, inference services, distributed systems, and production settings.

• Assess emerging models, training techniques, inference frameworks, and machine learning infrastructure.

• Collaborate with AI engineers working on agentic systems to provide model, evaluation, and training infrastructure.


⛳️ Requirements

• U.S. Citizenship is required.

• Over 5 years of experience in building production machine learning systems, ML infrastructure, distributed systems, or similar technical environments.

• Extensive experience in designing, constructing, and managing production ML infrastructure or ML platforms.

• Proven experience in building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models or other large-scale deep learning models.

• Familiarity with LLM fine-tuning and post-training workflows, including supervised fine-tuning, LoRA/QLoRA or other parameter-efficient methods, and preference optimization.

• Strong understanding of the contemporary LLM lifecycle, encompassing data preparation, training, evaluation, model artifacts, deployment, inference, monitoring, and iteration.

• Experience in building reproducible ML pipelines with dataset versioning, experiment tracking, model versioning, and automated evaluation.

• Proven competency in building and operating production GPU infrastructure across AWS, GCP, Azure, or dedicated GPU providers, including training and/or inference workloads.

• Strong knowledge of distributed systems and computationally intensive ML workloads at scale.

• Proficient programming skills in Python and experience in producing production-quality software.

• Extensive experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure.

• Experience in designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.

• Robust understanding of observability and operational reliability for production ML systems.

• Comfortable with debugging failures across training code, datasets, models, GPUs, distributed systems, and cloud infrastructure.

• Ability to transition smoothly between ML experimentation and infrastructure engineering.

• Willingness to work in a fast-paced and evolving field.

• Current possession of a U.S. security clearance, or the capacity to obtain one with sponsorship (preferred).

• Experience in or exposure to a startup or entrepreneurial setting (preferred).

• Background in building secure code execution environments or sandboxes for AI agents (preferred).

• Familiarity with multi-agent architectures, agent-to-agent communication, or distributed agent execution (preferred).

• Experience with fine-tuning, post-training, reinforcement learning, or synthetic data generation (preferred).

• Experience in constructing AI observability, tracing, and debugging infrastructure (preferred).

• Experience optimizing inference latency, throughput, GPU utilization, or model-serving expenses (preferred).

• Knowledge of AI security, adversarial testing, or securing agentic systems (preferred).

• Experience working in government, defense, or other mission-critical environments (preferred).


🏝️ Benefits

• Up to 25% travel, including occasional trips to Pittsburgh, PA and Arlington, VA offices for team collaboration, planning sessions, and in-person meetings.

• Security clearance sponsorship is available for candidates who can obtain a U.S. security clearance.

• Equal Opportunity Employer.

People also viewed

ClickUp20 hours ago

Machine Learning Engineer, Ranking & Retrieval

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $250k/year
ApplyView job
U-Glow21 hours ago

Intern – LLM/RAG, Machine Learning Engineer

US flagUnited States OnlyInternshipMachine Learning Engineer
ApplyView job
Neon22 hours ago

Staff Machine Learning Engineer

BR flagBrazil OnlyFull-timeMachine Learning Engineer
ApplyView job
Imnoo1 day ago

Machine Learning, Computer Graphics Engineer – Junior/Mid/Senior

CH flagSwitzerland OnlyFull-timeMachine Learning Engineer
ApplyView job
OmegaHires1 day ago

AI, Machine Learning Roles

US flagUnited States OnlyFreelanceMachine Learning Engineer$87 – $117/hour
ApplyView job
24-MAG1 day ago

Machine Learning Engineer

US flagNew York OnlyPart-timeMachine Learning Engineer$60 – $80/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers