Remotery

AI Infrastructure Engineer III

Posted Aug 4

This is a fully remote position, open to applicants in Egypt.

📋 Description

• Design, implement, and manage enterprise AI/ML platforms.

• Develop self-service solutions for Data Scientists and ML Engineers.

• Deploy and maintain Kubeflow, MLflow, KServe, Ray, or comparable AI platforms.

• Create infrastructure that supports model training, experimentation, feature engineering, and inference.

• Construct a highly available and scalable model serving infrastructure.

• Design and manage GPU clusters to handle large-scale AI workloads.

• Optimize GPU scheduling, utilization, sharing, autoscaling, and resource allocation.

• Deploy and oversee NVIDIA GPU Operator and GPU-enabled Kubernetes environments.

• Enhance distributed GPU training performance across multi-node clusters.

• Diagnose and resolve performance bottlenecks in AI infrastructure.

• Develop CI/CD pipelines for ML workloads.

• Automate AI infrastructure provisioning through Infrastructure as Code.

• Implement monitoring and observability for GPU utilization, model serving, training jobs, and inference latency.

• Collaborate with Data Science teams to enhance platform usability, performance, and reliability.


⛳️ Requirements

• 4–6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.

• Extensive hands-on experience with Kubernetes.

• Familiarity with Kubeflow, MLflow, or similar ML platform technologies.

• Experience in managing GPU infrastructure for AI workloads.

• In-depth understanding of NVIDIA GPU technologies, CUDA fundamentals, and GPU optimization.

• Experience supporting distributed training workloads.

• Proficiency with model serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.

• Familiarity with AWS, GCP, OCI, or Azure AI platforms.

• Experience in automating infrastructure using Terraform, Helm, GitOps, or Ansible.

• Strong scripting or programming abilities in Python, Bash, or Go.

• Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar observability platforms.

• Preferred: experience with PyTorch, TensorFlow, Hugging Face, or JAX.

• Preferred: experience with distributed training frameworks such as Ray, DeepSpeed, Horovod, or NCCL.

• Preferred: experience with vector databases, LLM infrastructure, RAG architectures, or GenAI platforms.

• Preferred: experience managing inference platforms for large language models.

• Preferred: experience supporting AI research or Data Science teams in production settings.

• Preferred: contributions to Cloud Native, Kubernetes, AI, or ML open-source communities.

• Cloud, Kubernetes, NVIDIA, or AI/ML certifications are a plus.


🏝️ Benefits

• Competitive compensation.

• Top-tier health insurance.

• Enabling culture.

• Responsibility and trust.

• Freedom and autonomy in the role.

• Fun and dynamic workplace.

• Opportunity to work alongside leading AI professionals.

• Inclusive and empowering workplace culture.

People also viewed

adconova GmbH2 days ago

Senior Cloud & AI Infrastructure Engineer

DE flagGermany OnlyFull-timeInfrastructure Engineer€60k – €100k/year
ApplyView job
Teleperformance3 days ago

Senior Systems Engineer – IT Infrastructure Engineer

HR flagCroatia OnlyFull-timeInfrastructure Engineer
ApplyView job
Trilon Group3 days ago

AWS Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$90k – $110k/year
ApplyView job
Carbon603 days ago

Principal Infrastructure Architect

CA flagCanada OnlyFull-timeInfrastructure EngineerC$180k – C$220k/year
ApplyView job
fal3 days ago

Senior/Staff Kubernetes Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$180k – $250k/year
ApplyView job
LatamCent4 days ago

Lead Security and Infrastructure Engineer

US flagFlorida OnlyFull-timeInfrastructure Engineer$170k – $210k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers