Remotery

Senior AI Tools Engineer, SRE Operations

Posted Aug 5

This is a fully remote position, open to applicants in California.

📋 Description

• Develop and deploy advanced AI-driven tools and products that enhance the operation and optimization of the global GeForce NOW service.

• Convert production data streams, which encompass signals, metrics, and logs, into actionable insights.

• Automate root cause analysis for incidents while forecasting future service trends and patterns.

• Create and implement AI/ML tools aimed at identifying the root causes of incidents and recognizing operational trends.

• Spearhead the development of LLM- and agent-based systems to boost operational efficiency.

• Establish and uphold data management practices and pipelines for extensive model-development datasets.

• Take ownership of and improve LLM-based pipelines.

• Integrate advancements in LLM technology into product development.

• Act as a subject matter expert on AI frameworks, advising on platforms, toolsets, and architectures for long-term product viability.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Statistics, Engineering, or a related field, or equivalent professional experience.

• Over 5 years of relevant experience.

• Proficient in Python programming.

• Familiarity with Go or other systems programming languages is advantageous.

• Practical experience in building, optimizing, and deploying AI tools.

• Deep understanding of AI advancements and LLM-based platforms.

• Hands-on experience with Kubernetes and AWS cloud platforms.

• Ability to discern significant AI advancements from irrelevant information in technical decision-making.

• Expertise in automation and large-scale data processing pipelines.

• Experience with monitoring and visualization tools, such as Grafana.

• Exceptional skills in handling, transforming, and managing data sources and pipelines.

• Understanding of SRE principles and management of production environments is preferred.

• Familiarity with LLM enhancement pipelines and recent developments in LLM training is preferred.

• Knowledge of LLMs and AI models, including the ability to recommend sustainable platforms and methodologies.


🏝️ Benefits

• Equity.

• Comprehensive benefits.

• Competitive salary package.

• Inclusive and supportive workplace culture.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers