
Platform Engineer
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in Canada, +1 more country.
• Create and establish scalable infrastructure for LLM and GenAI tasks within multi-GPU environments.
• Conduct GPU profiling, benchmarking, and performance enhancements for distributed training tasks.
• Oversee and schedule compute-heavy jobs utilizing Slurm-based clusters and OpenShift/Kubernetes setups.
• Activate and enhance the NVIDIA GPU stack.
• Collaborate with architects, data scientists, MLOps, application teams, and AI solution teams.
• Implement models in both research and production settings.
• Develop and maintain GenAI pipelines, encompassing fine-tuning, RAG, multi-modal inferencing, and LLMOps.
• Create reusable infrastructure templates using Terraform and Helm.
• Contribute to internal PoCs and workshops while supporting client-facing delivery engagements.
• Design automation software to enhance functionality, reliability, availability, and manageability of applications and cloud platforms.
• Promote the adoption of Infrastructure as Code principles.
• Design and construct self-service, self-healing, synthetic monitoring, and alerting platforms and tools.
• Automate development and testing workflows through CI/CD pipelines utilizing Git, Jenkins, SonarQube, Artifactory, and Docker.
• Develop container hosting solutions using Kubernetes.
• Introduce cloud technologies and tools to generate business value.
• Lead technical discussions with clients on architecture design and troubleshooting, proactively offering solutions.
• Guide and mentor senior resources and team leads.
• Over 10 years of relevant experience.
• Extensive experience with Slurm and distributed training setups.
• Practical knowledge of Red Hat OpenShift and/or Kubernetes.
• In-depth understanding of the NVIDIA GPU ecosystem: CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT.
• Solid background in Linux systems, performance tuning, and multi-GPU optimization.
• Experience in deploying GenAI tasks, including LLM fine-tuning, RAG pipelines, and multi-modal systems.
• Familiarity with Infrastructure-as-Code tools like Terraform and Ansible.
• Proficient with cloud GPU environments: GCP, Azure, AWS, OCI, and/or on-premises GPU clusters.
• Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers.
• Knowledge of LLMOps frameworks and integration with MLOps.
• Familiar with vector databases and retrieval systems for RAG architecture.
• Comfortable in client-facing roles and collaborating with AI solution teams.
• Experience in the healthcare domain is a plus, including FHIR R4, HL7 v2, SMART on FHIR, EHR integration, HIPAA, CDS Hooks, clinical workflows, and clinical decision support systems.
• Advance your skills and explore your potential through innovative technology projects.
• Work within a research-driven organization that has filed over 60 patents.
• Gain exposure to AI, ML, data, and cloud technologies.
• Collaborate with Fortune 500 companies.
• Enjoy opportunities for learning, growth, and global colleague interactions.
• Embrace a hybrid work culture.
Encompass Corporation
Shippit
Group O
ServiceTitan
Get handpicked remote jobs straight to your inbox weekly.