Software Engineer, Inference

Posted 2 days ago

This is a fully remote position, open to applicants in Europe.

📋 Description

• Take ownership of how Luma's models are served through the integration of the inference engine, scaling deployments, and optimizing GPU fleet utilization.

• Incorporate new model architectures into the inference engine.

• Work collaboratively with research, engineering, and infrastructure teams to enhance model efficiency and deployment processes.

• Develop internal tools to measure, profile, and monitor inference jobs and workflows.

• Automate, test, and maintain inference services to ensure uptime and reliability.

• Oversee and optimize inference workloads across various clusters and hardware providers.

• Scale deployments to operate across thousands of machines.

• Create scheduling systems that maximize GPU resource utilization while adhering to service level objectives (SLOs).

• Maintain continuous integration and continuous deployment (CI/CD) for model checkpoints and software development kits (SDKs).

• Acquire knowledge of the inference stack and troubleshoot reliability or utilization issues within the first 30 days.

• Integrate a model or implement tooling/scheduling enhancements during the period of 30 to 60 days.

• Strengthen deployment pipelines and scheduling across clusters and providers in the timeframe of 60 to 90 days.


⛳️ Requirements

• Proficient in Python and system architecture.

• Proven experience in deploying models using PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar technologies.

• Familiarity with queues, scheduling, traffic management, and fleet administration at scale.

• Competence in Linux, Docker, and Kubernetes.

• Experience in orchestration, deployment, and scheduling practices.

• Knowledge of Redis and S3-compatible storage solutions.

• Preferred: familiarity with modern networking stacks including RDMA (RoCE, InfiniBand, NVLink).

• Preferred: experience with high-performance large-scale machine learning systems involving 100 or more GPUs.

• Preferred: knowledge of CUDA, and FFmpeg or multimedia processing.


🏝️ Benefits

• Commitment to being an equal opportunity employer.

• Participation in a voluntary diversity and inclusion survey; opting out will not impact the job application.

• Flexible remote work arrangement.

People also viewed

IQ Plus AG13 hours ago

Software Architect, Microservice Architecture

CH flagSwitzerland OnlyFull-timeFull-stack Engineer
ApplyView job
Clipbook16 hours ago

Founding Senior Software Engineer – Europe

EuropeFull-timeFull-stack Engineer
ApplyView job
GFT Technologies20 hours ago

Senior AI Software Engineer

CR flagCosta Rica OnlyFull-timeFull-stack Engineer
ApplyView job
Anyone AI23 hours ago

Senior Software Engineer – Open Source, SWE-Bench Evaluation

AR flagArgentina OnlyPart-timeFull-stack Engineer$65/hour
ApplyView job
S + S Regeltechnik GmbH1 day ago

Senior Full-Stack Developer

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
eCom Solutions Inc1 day ago

Senior Full-Stack Engineer, React/Node.js

AM flagArmenia OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers