
AI Infrastructure Engineer III
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Egypt.
• Design, implement, and manage enterprise AI/ML platforms.
• Develop self-service solutions for Data Scientists and ML Engineers.
• Deploy and maintain Kubeflow, MLflow, KServe, Ray, or comparable AI platforms.
• Create infrastructure that supports model training, experimentation, feature engineering, and inference.
• Construct a highly available and scalable model serving infrastructure.
• Design and manage GPU clusters to handle large-scale AI workloads.
• Optimize GPU scheduling, utilization, sharing, autoscaling, and resource allocation.
• Deploy and oversee NVIDIA GPU Operator and GPU-enabled Kubernetes environments.
• Enhance distributed GPU training performance across multi-node clusters.
• Diagnose and resolve performance bottlenecks in AI infrastructure.
• Develop CI/CD pipelines for ML workloads.
• Automate AI infrastructure provisioning through Infrastructure as Code.
• Implement monitoring and observability for GPU utilization, model serving, training jobs, and inference latency.
• Collaborate with Data Science teams to enhance platform usability, performance, and reliability.
• 4–6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.
• Extensive hands-on experience with Kubernetes.
• Familiarity with Kubeflow, MLflow, or similar ML platform technologies.
• Experience in managing GPU infrastructure for AI workloads.
• In-depth understanding of NVIDIA GPU technologies, CUDA fundamentals, and GPU optimization.
• Experience supporting distributed training workloads.
• Proficiency with model serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.
• Familiarity with AWS, GCP, OCI, or Azure AI platforms.
• Experience in automating infrastructure using Terraform, Helm, GitOps, or Ansible.
• Strong scripting or programming abilities in Python, Bash, or Go.
• Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar observability platforms.
• Preferred: experience with PyTorch, TensorFlow, Hugging Face, or JAX.
• Preferred: experience with distributed training frameworks such as Ray, DeepSpeed, Horovod, or NCCL.
• Preferred: experience with vector databases, LLM infrastructure, RAG architectures, or GenAI platforms.
• Preferred: experience managing inference platforms for large language models.
• Preferred: experience supporting AI research or Data Science teams in production settings.
• Preferred: contributions to Cloud Native, Kubernetes, AI, or ML open-source communities.
• Cloud, Kubernetes, NVIDIA, or AI/ML certifications are a plus.
• Competitive compensation.
• Top-tier health insurance.
• Enabling culture.
• Responsibility and trust.
• Freedom and autonomy in the role.
• Fun and dynamic workplace.
• Opportunity to work alongside leading AI professionals.
• Inclusive and empowering workplace culture.
adconova GmbH
Teleperformance
Trilon Group
Carbon60
Get handpicked remote jobs straight to your inbox weekly.