Python Inference Engineer

Posted Sep 9

This is a fully remote position, open to applicants in Cyprus, +2 more countries.

📋 Description

• Develop and enhance the inference layer of the Gcore Inference platform.

• Integrate and manage inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.

• Deploy new language and multimodal models into production.

• Optimize inference latency, throughput, memory usage, GPU utilization, and cost efficiency.

• Troubleshoot performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.

• Collaborate with platform, infrastructure, product, and customer-facing teams to transform inference enhancements into dependable product features.

• Contribute enhancements to open-source inference projects when relevant.


⛳️ Requirements

• Over 5 years of experience in writing reliable, well-tested production code.

• Proficient in Python with experience in designing production systems.

• Practical experience with PyTorch and deploying machine learning models.

• Familiarity with Linux, Docker, and Kubernetes.

• Experience in at least one relevant area: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling.

• Capability to debug intricate issues across software, infrastructure, and hardware.

• Strong understanding of developer experience.

• Genuine enthusiasm for inference engineering and a desire to learn.

• Effective communication and collaboration skills.

• Nice to have: experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a similar inference framework.

• Nice to have: experience managing GPU workloads in a production environment.

• Nice to have: knowledge of quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving.

• Nice to have: experience with CUDA, Triton, TensorRT, or other GPU programming tools.

• Nice to have: experience in profiling and enhancing model latency, throughput, memory usage, or GPU utilization.

• Nice to have: experience with distributed inference, multi-GPU systems, scheduling, or autoscaling.

• Nice to have: contributions to open-source ML, inference, or infrastructure projects.


🏝️ Benefits

• Competitive compensation.

• Flexible working hours and hybrid or remote options, based on your role.

• Work from any location in the world for up to 45 days each year.

• Private medical insurance for you and your family.*

• Additional paid vacation and sick leave days.*

• Support for significant life moments and celebrations.

• Language courses to enhance your connections and growth.

• Modern, inviting offices stocked with snacks, beverages, and entertainment.*

• Team sports and social activities.*

People also viewed

MWDN1 day ago

Senior Python Backend Developer

UA flagUkraine OnlyFull-timeBackend Engineer
ApplyView job
Oscilar1 day ago

Senior Backend Engineer – Java

IN flagIndia OnlyFull-timeBackend Engineer₹6.6M – ₹9.7M/year
ApplyView job
Veradigm®1 day ago

Principal Database Engineer – Platform Engineer

US flagUnited States, +1 more countryFull-timeBackend Engineer$130.2k – $189.5k/year
ApplyView job
Shuru1 day ago

Lead Java Engineer

IN flagIndia OnlyFull-timeBackend Engineer
ApplyView job
FCamara Consulting & Training1 day ago

Senior Backend Developer – Kotlin

BR flagBrazil OnlyFull-timeBackend Engineer
ApplyView job
Sedgwick1 day ago

Power Platform Architect – AI

US flagFlorida OnlyFull-timeBackend Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers