
Software Engineer, Inference
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
• Take ownership of how Luma's models are served through the integration of the inference engine, scaling deployments, and optimizing GPU fleet utilization.
• Incorporate new model architectures into the inference engine.
• Work collaboratively with research, engineering, and infrastructure teams to enhance model efficiency and deployment processes.
• Develop internal tools to measure, profile, and monitor inference jobs and workflows.
• Automate, test, and maintain inference services to ensure uptime and reliability.
• Oversee and optimize inference workloads across various clusters and hardware providers.
• Scale deployments to operate across thousands of machines.
• Create scheduling systems that maximize GPU resource utilization while adhering to service level objectives (SLOs).
• Maintain continuous integration and continuous deployment (CI/CD) for model checkpoints and software development kits (SDKs).
• Acquire knowledge of the inference stack and troubleshoot reliability or utilization issues within the first 30 days.
• Integrate a model or implement tooling/scheduling enhancements during the period of 30 to 60 days.
• Strengthen deployment pipelines and scheduling across clusters and providers in the timeframe of 60 to 90 days.
• Proficient in Python and system architecture.
• Proven experience in deploying models using PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar technologies.
• Familiarity with queues, scheduling, traffic management, and fleet administration at scale.
• Competence in Linux, Docker, and Kubernetes.
• Experience in orchestration, deployment, and scheduling practices.
• Knowledge of Redis and S3-compatible storage solutions.
• Preferred: familiarity with modern networking stacks including RDMA (RoCE, InfiniBand, NVLink).
• Preferred: experience with high-performance large-scale machine learning systems involving 100 or more GPUs.
• Preferred: knowledge of CUDA, and FFmpeg or multimedia processing.
• Commitment to being an equal opportunity employer.
• Participation in a voluntary diversity and inclusion survey; opting out will not impact the job application.
• Flexible remote work arrangement.
IQ Plus AG
Clipbook
GFT Technologies
Anyone AI
Get handpicked remote jobs straight to your inbox weekly.