
Platform Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Canada, +1 more country.
β’ Create and establish scalable infrastructure for LLM and GenAI tasks within multi-GPU environments.
β’ Conduct GPU profiling, benchmarking, and performance enhancements for distributed training tasks.
β’ Oversee and schedule compute-heavy jobs utilizing Slurm-based clusters and OpenShift/Kubernetes setups.
β’ Activate and enhance the NVIDIA GPU stack.
β’ Collaborate with architects, data scientists, MLOps, application teams, and AI solution teams.
β’ Implement models in both research and production settings.
β’ Develop and maintain GenAI pipelines, encompassing fine-tuning, RAG, multi-modal inferencing, and LLMOps.
β’ Create reusable infrastructure templates using Terraform and Helm.
β’ Contribute to internal PoCs and workshops while supporting client-facing delivery engagements.
β’ Design automation software to enhance functionality, reliability, availability, and manageability of applications and cloud platforms.
β’ Promote the adoption of Infrastructure as Code principles.
β’ Design and construct self-service, self-healing, synthetic monitoring, and alerting platforms and tools.
β’ Automate development and testing workflows through CI/CD pipelines utilizing Git, Jenkins, SonarQube, Artifactory, and Docker.
β’ Develop container hosting solutions using Kubernetes.
β’ Introduce cloud technologies and tools to generate business value.
β’ Lead technical discussions with clients on architecture design and troubleshooting, proactively offering solutions.
β’ Guide and mentor senior resources and team leads.
β’ Over 10 years of relevant experience.
β’ Extensive experience with Slurm and distributed training setups.
β’ Practical knowledge of Red Hat OpenShift and/or Kubernetes.
β’ In-depth understanding of the NVIDIA GPU ecosystem: CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT.
β’ Solid background in Linux systems, performance tuning, and multi-GPU optimization.
β’ Experience in deploying GenAI tasks, including LLM fine-tuning, RAG pipelines, and multi-modal systems.
β’ Familiarity with Infrastructure-as-Code tools like Terraform and Ansible.
β’ Proficient with cloud GPU environments: GCP, Azure, AWS, OCI, and/or on-premises GPU clusters.
β’ Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers.
β’ Knowledge of LLMOps frameworks and integration with MLOps.
β’ Familiar with vector databases and retrieval systems for RAG architecture.
β’ Comfortable in client-facing roles and collaborating with AI solution teams.
β’ Experience in the healthcare domain is a plus, including FHIR R4, HL7 v2, SMART on FHIR, EHR integration, HIPAA, CDS Hooks, clinical workflows, and clinical decision support systems.
β’ Advance your skills and explore your potential through innovative technology projects.
β’ Work within a research-driven organization that has filed over 60 patents.
β’ Gain exposure to AI, ML, data, and cloud technologies.
β’ Collaborate with Fortune 500 companies.
β’ Enjoy opportunities for learning, growth, and global colleague interactions.
β’ Embrace a hybrid work culture.
Resend
Resend
Improvix Technologies
Improvix Technologies
Get handpicked remote jobs straight to your inbox weekly.