
Senior AI Deployment Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Texas.
• Collaborate with internationally distributed teams to tackle technical challenges and enhance processes.
• Design, implement, maintain, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.
• Work closely with stakeholders to gather and refine technical specifications.
• Improve system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Engage in code reviews to uphold high quality standards.
• Keep abreast of industry trends and best practices in cloud operations and development.
• Create and implement AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.
• Enable knowledge transfer to clients during delivery phases.
• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Background in high-performance data center processing, networking, and storage.
• Familiarity with Golang and working knowledge of additional programming languages (Python, JavaScript).
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Outstanding problem-solving and debugging abilities across networking and storage (both hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven capability to lead technical initiatives and collaborate effectively with diverse teams.
• Comfortable making independent decisions when interacting directly with clients, often with minimal oversight.
• Excellent proficiency in written and spoken English.
• Strong customer-facing communication skills.
• A dedication to innovation, ongoing learning, and delivering high-quality outcomes.
• Willingness to travel up to 25% if necessary, including internationally.
• Nice to have extensive experience in network and/or storage architecture.
• Experience in high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Active involvement in the open source community, including upstream contributions and conference presentations.
• Previous experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.
• Collaborate with a recognized Silicon Valley leader in the cloud infrastructure sector;
• Work alongside exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies;
• Take part in cutting-edge, open-source innovation;
• Flourish in the dynamic environment of a young company that values openness, collaboration, risk-taking, and continuous growth;
• Access to professional development and training;
• Opportunities to attend conferences and participate in working groups;
• Enjoy company outings, happy hours, hackathons, and tech talks;
• Receive a competitive salary package accompanied by a robust benefits plan.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.