
Senior Machine Learning Engineer
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Pennsylvania, +1 more state.
• Develop infrastructure that supports the development, evaluation, deployment, and ongoing enhancement of language models and AI systems.
• Take ownership of LLMOps, fine-tuning infrastructure, model evaluation, dataset pipelines, experiment management, model serving, and production observability.
• Create platforms that allow AI engineers and researchers to conduct experiments swiftly while ensuring reproducibility, scalability, and reliability.
• Assist in the complete model lifecycle from dataset creation and experimentation to training, evaluation, deployment, monitoring, and iteration.
• Design and construct LLMOps infrastructure tailored for production language models.
• Establish scalable training and fine-tuning infrastructure for both commercial and open-weight language models.
• Develop pipelines for supervised fine-tuning, parameter-efficient fine-tuning, preference optimization, and various post-training methodologies.
• Construct distributed training and GPU-accelerated machine learning infrastructure.
• Create data pipelines for training, fine-tuning, evaluation, and synthetic data generation.
• Implement dataset versioning, lineage, quality validation, transformation, and reproducible experimentation systems.
• Develop infrastructure for experiment management to compare models, datasets, hyperparameters, prompts, and training methods.
• Build automated model evaluation pipelines and checks for production readiness.
• Design model registries, artifact management, versioning, and promotion workflows.
• Construct and manage scalable model-serving and inference infrastructure.
• Create abstractions that support multiple models and inference providers.
• Implement observability for training and inference, covering metrics, tracing, logging, resource utilization, model quality, latency, throughput, and cost.
• Optimize workloads for GPU efficiency, throughput, latency, reliability, and infrastructure expenses.
• Establish automated workflows for deployment, rollback, canarying, and production validation.
• Investigate failures across data pipelines, training jobs, inference services, distributed systems, and production settings.
• Assess emerging models, training techniques, inference frameworks, and machine learning infrastructure.
• Collaborate with AI engineers working on agentic systems to provide model, evaluation, and training infrastructure.
• U.S. Citizenship is required.
• Over 5 years of experience in building production machine learning systems, ML infrastructure, distributed systems, or similar technical environments.
• Extensive experience in designing, constructing, and managing production ML infrastructure or ML platforms.
• Proven experience in building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models or other large-scale deep learning models.
• Familiarity with LLM fine-tuning and post-training workflows, including supervised fine-tuning, LoRA/QLoRA or other parameter-efficient methods, and preference optimization.
• Strong understanding of the contemporary LLM lifecycle, encompassing data preparation, training, evaluation, model artifacts, deployment, inference, monitoring, and iteration.
• Experience in building reproducible ML pipelines with dataset versioning, experiment tracking, model versioning, and automated evaluation.
• Proven competency in building and operating production GPU infrastructure across AWS, GCP, Azure, or dedicated GPU providers, including training and/or inference workloads.
• Strong knowledge of distributed systems and computationally intensive ML workloads at scale.
• Proficient programming skills in Python and experience in producing production-quality software.
• Extensive experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure.
• Experience in designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.
• Robust understanding of observability and operational reliability for production ML systems.
• Comfortable with debugging failures across training code, datasets, models, GPUs, distributed systems, and cloud infrastructure.
• Ability to transition smoothly between ML experimentation and infrastructure engineering.
• Willingness to work in a fast-paced and evolving field.
• Current possession of a U.S. security clearance, or the capacity to obtain one with sponsorship (preferred).
• Experience in or exposure to a startup or entrepreneurial setting (preferred).
• Background in building secure code execution environments or sandboxes for AI agents (preferred).
• Familiarity with multi-agent architectures, agent-to-agent communication, or distributed agent execution (preferred).
• Experience with fine-tuning, post-training, reinforcement learning, or synthetic data generation (preferred).
• Experience in constructing AI observability, tracing, and debugging infrastructure (preferred).
• Experience optimizing inference latency, throughput, GPU utilization, or model-serving expenses (preferred).
• Knowledge of AI security, adversarial testing, or securing agentic systems (preferred).
• Experience working in government, defense, or other mission-critical environments (preferred).
• Up to 25% travel, including occasional trips to Pittsburgh, PA and Arlington, VA offices for team collaboration, planning sessions, and in-person meetings.
• Security clearance sponsorship is available for candidates who can obtain a U.S. security clearance.
• Equal Opportunity Employer.
ClickUp
U-Glow
Neon
Imnoo
Get handpicked remote jobs straight to your inbox weekly.