Remotery

Senior AI Deployment Engineer

Posted Jul 29

This is a fully remote position, open to applicants in Texas.

📋 Description

• Collaborate with internationally distributed teams to tackle technical challenges and enhance processes.

• Design, implement, maintain, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.

• Work closely with stakeholders to gather and refine technical specifications.

• Improve system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Engage in code reviews to uphold high quality standards.

• Keep abreast of industry trends and best practices in cloud operations and development.

• Create and implement AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.

• Enable knowledge transfer to clients during delivery phases.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies, including Kubernetes and/or OpenStack.

• Background in high-performance data center processing, networking, and storage.

• Familiarity with Golang and working knowledge of additional programming languages (Python, JavaScript).

• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Outstanding problem-solving and debugging abilities across networking and storage (both hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.

• Proven capability to lead technical initiatives and collaborate effectively with diverse teams.

• Comfortable making independent decisions when interacting directly with clients, often with minimal oversight.

• Excellent proficiency in written and spoken English.

• Strong customer-facing communication skills.

• A dedication to innovation, ongoing learning, and delivering high-quality outcomes.

• Willingness to travel up to 25% if necessary, including internationally.

• Nice to have extensive experience in network and/or storage architecture.

• Experience in high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.

• Active involvement in the open source community, including upstream contributions and conference presentations.

• Previous experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.


🏝️ Benefits

• Collaborate with a recognized Silicon Valley leader in the cloud infrastructure sector;

• Work alongside exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies;

• Take part in cutting-edge, open-source innovation;

• Flourish in the dynamic environment of a young company that values openness, collaboration, risk-taking, and continuous growth;

• Access to professional development and training;

• Opportunities to attend conferences and participate in working groups;

• Enjoy company outings, happy hours, hackathons, and tech talks;

• Receive a competitive salary package accompanied by a robust benefits plan.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers