
Python Inference Engineer
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Cyprus, +2 more countries.
• Develop and enhance the inference layer of the Gcore Inference platform.
• Integrate and manage inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
• Deploy new language and multimodal models into production.
• Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency.
• Troubleshoot performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
• Collaborate with platform, infrastructure, product, and customer-facing teams to transform inference enhancements into dependable product features.
• Contribute enhancements to open-source inference projects when relevant.
• Over 5 years of experience in writing reliable, well-tested production code.
• Proficient in Python with experience in designing production systems.
• Practical experience with PyTorch and deploying machine learning models.
• Familiarity with Linux, Docker, and Kubernetes.
• Experience in at least one relevant area: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.
• Capability to debug intricate issues across software, infrastructure, and hardware.
• Strong understanding of developer experience.
• Genuine enthusiasm for inference engineering and a desire to learn.
• Effective communication and collaboration skills.
• Nice to have: experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a similar inference framework.
• Nice to have: experience managing GPU workloads in a production environment.
• Nice to have: knowledge of quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving.
• Nice to have: experience with CUDA, Triton, TensorRT, or other GPU programming tools.
• Nice to have: experience in profiling and enhancing model latency, throughput, memory usage, or GPU utilization.
• Nice to have: experience with distributed inference, multi-GPU systems, scheduling, or autoscaling.
• Nice to have: contributions to open-source ML, inference, or infrastructure projects.
• Competitive compensation.
• Flexible working hours and hybrid or remote options, based on your role.
• Work from any location in the world for up to 45 days each year.
• Private medical insurance for you and your family.*
• Additional paid vacation and sick leave days.*
• Support for significant life moments and celebrations.
• Language courses to enhance your connections and growth.
• Modern, inviting offices stocked with snacks, beverages, and entertainment.*
• Team sports and social activities.*
Oscilar
Veradigm®
Get handpicked remote jobs straight to your inbox weekly.