Python Inference Engineer

Posted Sep 9

This is a fully remote position, open to applicants in Cyprus, +2 more countries.

📋 Description

• Develop and enhance the inference layer of the Gcore Inference platform

• Integrate and manage inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM

• Deploy new language and multimodal models into production environments

• Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency

• Troubleshoot performance and reliability challenges across model code, inference frameworks, GPU execution, networking, and Kubernetes

• Collaborate with platform, infrastructure, product, and customer-facing teams to transform inference enhancements into dependable product features

• Contribute enhancements to open-source inference projects when suitable


⛳️ Requirements

• Over 5 years of experience in writing reliable, thoroughly tested production code

• Proficient in Python with a strong background in designing production systems

• Practical experience with PyTorch and deploying machine learning models

• Familiarity with Linux, Docker, and Kubernetes

• Experience in at least one relevant field: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling

• Capability to debug intricate issues across software, infrastructure, and hardware

• Strong emphasis on developer experience

• Genuine enthusiasm for inference engineering

• Eagerness to learn about inference engines

• Effective communication and teamwork abilities

• Nice to have: Familiarity with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a comparable inference framework

• Nice to have: Experience managing GPU workloads in a production setting

• Nice to have: Understanding of inference techniques such as quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving

• Nice to have: Experience with CUDA, Triton, TensorRT, or other GPU programming tools

• Nice to have: Experience profiling and enhancing model latency, throughput, memory usage, or GPU utilization

• Nice to have: Knowledge of distributed inference, multi-GPU systems, scheduling, or autoscaling

• Nice to have: Contributions to open-source ML, inference, or infrastructure projects


🏝️ Benefits

• Competitive salary

• Flexible working hours along with hybrid or remote options, depending on your role

• Ability to work from anywhere in the world for up to 45 days each year

• Private medical insurance for you and your family*

• Additional paid vacation and sick leave days*

• Support for significant life events and celebrations

• Language courses to aid in connection and personal growth

• Modern, inviting offices equipped with snacks, drinks, and entertainment*

• Opportunities for team sports and social activities*

People also viewed

MWDN1 day ago

Senior Python Backend Developer

UA flagUkraine OnlyFull-timeBackend Engineer
ApplyView job
Oscilar1 day ago

Senior Backend Engineer – Java

IN flagIndia OnlyFull-timeBackend Engineer₹6.6M – ₹9.7M/year
ApplyView job
Veradigm®1 day ago

Principal Database Engineer – Platform Engineer

US flagUnited States, +1 more countryFull-timeBackend Engineer$130.2k – $189.5k/year
ApplyView job
Shuru1 day ago

Lead Java Engineer

IN flagIndia OnlyFull-timeBackend Engineer
ApplyView job
FCamara Consulting & Training1 day ago

Senior Backend Developer – Kotlin

BR flagBrazil OnlyFull-timeBackend Engineer
ApplyView job
Sedgwick1 day ago

Power Platform Architect – AI

US flagFlorida OnlyFull-timeBackend Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers