
Senior MLOps Engineer, LLMOps
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Develop reusable CI/CD workflows for model training, evaluation, and deployment — incorporating Langfuse, GitHub Actions, and experiment tracking, among others.
• Automate processes for model versioning, approval workflows, and compliance checks across various environments.
• Construct a modular and scalable AI infrastructure stack — which includes vector databases, feature stores, model registries, and observability tools.
• Collaborate with engineering and data science teams to integrate AI models and agents into real-time applications and workflows.
• Regularly assess and incorporate cutting-edge AI tools (such as LangChain, LlamaIndex, vLLM, MLflow, BentoML, etc.).
• Promote AI reliability and governance, facilitating experimentation while ensuring compliance, security, and system uptime.
• Enhance the performance of AI/ML models.
• Guarantee data accuracy, consistency, and reliability, contributing to improved model training and inference.
• Deploy infrastructure that supports both offline and online evaluation of LLMs and agents — including regression testing, cost monitoring, and human-in-the-loop workflows.
• Equip researchers with tools to iterate rapidly by providing sandboxes, dashboards, and reproducible environments.
• Produce high-quality, maintainable software — primarily using Python.
• Possess a strong foundation in scalable infrastructure, including: Containerization and orchestration (e.g., Docker, Kubernetes).
• Experience with infrastructure-as-code and deployment (e.g., Terraform, CI/CD pipelines).
• Familiarity with monitoring and logging frameworks (e.g., Datadog, Prometheus, OpenTelemetry).
• Understand and apply ML Ops best practices, including: Model versioning and rollback strategies, automated evaluation and drift detection, and scalable model and agent serving infrastructure (e.g., vLLM, Triton, BentoML).
• Deploy and maintain LLM and agentic workflows in production, which includes: Monitoring cost, latency, and performance; capturing traces for analysis and debugging; and optimizing prompt/response flows with real-time data access.
• Exhibit strong ownership and pragmatism, balancing infrastructure elegance with iterative delivery and measurable impact.
• Participate in TRM’s equity plan.
NVIDIA
SentiLink
SentiLink
Leega
Get handpicked remote jobs straight to your inbox weekly.