
HPC / AI Network Technologist
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Massachusetts.
• Collect client networking requirements and create optimized solutions utilizing various vendors and technologies.
• Design, implement, and sustain high-throughput, low-latency network architectures for HPC and AI clusters.
• Deploy new HPC/AI solutions or enhance existing ones from the ground up.
• Stay updated on multiple Network Operating Systems, including Cumulus Linux and SONIC.
• Configure and fine-tune switches, routers, and firewalls for HPC and AI workloads.
• Monitor network performance and resolve connectivity, latency, and security issues.
• Implement and uphold network security protocols, including firewalls, intrusion detection systems, and secure access measures.
• Document network configurations, procedures, and troubleshooting processes.
• Provide consultation and assist in the daily management of clients’ research computing infrastructure.
• Maintain HPC and AI infrastructure within Linux-based environments.
• Lead technical discussions and represent Cambridge Computer to clients during engagements.
• Validate solution designs, ensure client requirements are met, and confirm technical feasibility and deployability.
• Define professional services deliverables and establish clear client expectations.
• Create documentation and facilitate client knowledge transfer.
• Advise on networking, storage, data protection, digital archiving, and other infrastructure technologies.
• Develop advanced knowledge and obtain certifications from vendors utilized in Cambridge Computer’s solution stack.
• A minimum of 5+ years in providing networking deployment services and/or cluster administration.
• An undergraduate degree in Computer Science, Computer Engineering, or a related science field is required.
• Comprehensive understanding of networking protocols, including TCP/IP, DNS, VLANs, and routing protocols.
• Practical experience with network equipment.
• Networking certifications from Cisco, NVIDIA Networking (Mellanox), Juniper, Arista, HPE Aruba, and other manufacturers are preferred.
• Strong knowledge of GPU-centric hardware/software and Linux system administration.
• Solid fundamentals in cluster design/management technologies, including Bright, Werewolf, and XCat.
• Experience with storage technologies and parallel filesystems, such as Lustre, GPFS, and BeeGFS.
• Proficient in networking and configuring Ethernet and InfiniBand network switches.
• Familiarity with HPC schedulers, including SLURM, UGE, and LSF.
• Knowledge of programming/libraries, including MPI and CUDA.
• Proficiency in scripting languages, including Bash and Python.
• Extensive knowledge of technology industry leaders and infrastructure vendors.
• Familiarity with containerized environments like Kubernetes and Docker along with their networking requirements.
• Capability to work remotely, independently, and without supervision.
• Willingness to travel approximately 50% of the time, including short day trips.
• Exceptional communication skills, ability to multitask, and a keen attention to detail.
• Effective problem-solving skills, organization, creativity, intellectual curiosity, ability to handle ambiguity, and adaptability to various personalities.
• Authorization to work in the United States on a full-time basis is required.
• Competitive salary.
• Multiple health insurance options.
• Medical FSA and Dependent Care FSA.
• Dental insurance.
• Vision insurance.
• 401(k) savings plan with employer matching.
• Employer-sponsored long-term disability.
• Paid holidays and PTO that increases with longevity at the company.
• Discounted health club membership.
• Convenient parking.
• Opportunities for growth.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.