
Python Inference Engineer
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Cyprus, +2 more countries.
• Develop and enhance the inference layer of the Gcore Inference platform
• Integrate and manage inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM
• Deploy new language and multimodal models into production environments
• Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency
• Troubleshoot performance and reliability challenges across model code, inference frameworks, GPU execution, networking, and Kubernetes
• Collaborate with platform, infrastructure, product, and customer-facing teams to transform inference enhancements into dependable product features
• Contribute enhancements to open-source inference projects when suitable
• Over 5 years of experience in writing reliable, thoroughly tested production code
• Proficient in Python with a strong background in designing production systems
• Practical experience with PyTorch and deploying machine learning models
• Familiarity with Linux, Docker, and Kubernetes
• Experience in at least one relevant field: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling
• Capability to debug intricate issues across software, infrastructure, and hardware
• Strong emphasis on developer experience
• Genuine enthusiasm for inference engineering
• Eagerness to learn about inference engines
• Effective communication and teamwork abilities
• Nice to have: Familiarity with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a comparable inference framework
• Nice to have: Experience managing GPU workloads in a production setting
• Nice to have: Understanding of inference techniques such as quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving
• Nice to have: Experience with CUDA, Triton, TensorRT, or other GPU programming tools
• Nice to have: Experience profiling and enhancing model latency, throughput, memory usage, or GPU utilization
• Nice to have: Knowledge of distributed inference, multi-GPU systems, scheduling, or autoscaling
• Nice to have: Contributions to open-source ML, inference, or infrastructure projects
• Competitive salary
• Flexible working hours along with hybrid or remote options, depending on your role
• Ability to work from anywhere in the world for up to 45 days each year
• Private medical insurance for you and your family*
• Additional paid vacation and sick leave days*
• Support for significant life events and celebrations
• Language courses to aid in connection and personal growth
• Modern, inviting offices equipped with snacks, drinks, and entertainment*
• Opportunities for team sports and social activities*
Oscilar
Veradigm®
Get handpicked remote jobs straight to your inbox weekly.