Remotery

Site Reliability Engineer – AI Enablement

Posted Jul 15

This is a fully remote position, open to applicants in United States.

📋 Description

• In your role as a Site Reliability Engineer on the Central AI team, you will assist Health Catalyst engineering teams in adopting AI in a responsible and effective manner.

• You will train and mentor engineering teams on how to seamlessly integrate AI into their development workflows, encompassing the use of AI-assisted coding tools, prompt engineering techniques, and agentic development methodologies.

• Your responsibilities will include evaluating AI system designs submitted through the Central AI intake process, offering actionable insights on integration patterns, reliability risks, observability issues, and adherence to AI governance standards.

• You will act as a technical resource within the organization’s AI governance framework, aiding teams in understanding and applying policies related to model access, data management, risk categorization, and practical responsible AI usage.

• Collaborating with engineering teams during the design and implementation phases of AI projects, you will provide hands-on assistance with LLM integration, RAG pipelines, agentic architectures, and AI service patterns.

• You will apply an SRE perspective to AI systems, guiding teams on observability, SLOs, failure modes, and the operational readiness of AI-powered services.

• As a subject matter expert, you will participate in incident calls to provide AI-specific guidance as necessary.

• You will contribute to the establishment of internal standards, reference architectures, and reusable patterns that facilitate the correct development of AI systems from the outset.

• Collaborating closely with product managers, data scientists, security, and compliance stakeholders, you will ensure that AI implementations meet organizational, regulatory, and clinical standards.

• You will maintain comprehensive documentation of AI architecture patterns, governance guidelines, and review decisions to support knowledge sharing and organizational learning.

• Staying abreast of the rapidly changing AI landscape—including LLM capabilities, agentic frameworks, AI safety research, and SRE practices for AI systems—you will bring pertinent insights back to the team.


⛳️ Requirements

• Demonstrated experience in designing and implementing AI systems in production, including LLM API integration (e.g., Azure AI Foundry, Anthropic Claude) and AI-native application patterns.

• Practical experience with at least one agentic or RAG framework (e.g., LangChain, LlamaIndex, Semantic Kernel, or similar).

• A solid SRE or platform engineering background, with a working understanding of observability, reliability principles, and operational best practices.

• The ability to assess AI architectures for reliability, security, governance alignment, and operational readiness, and to convey findings effectively to both technical and non-technical audiences.

• Experience in advising or empowering engineering teams through coaching, conducting reviews, or leading training on AI tools and best practices.

• Knowledge of AI governance concepts, including risk tiering, responsible AI principles, prompt safety, and access control for AI services.

• Experience with cloud infrastructure, particularly Azure or AWS, including managed AI/ML services.

• Familiarity with container-based architectures (Docker, Kubernetes) and CI/CD pipelines.

• Excellent written and verbal communication skills, capable of articulating complex AI concepts to audiences with varying technical backgrounds.

• A highly collaborative, self-directed individual motivated by the success of others in utilizing new technology.


🏝️ Benefits

• Flexible PTO

• Professional development stipend

• Remote-first work environment

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers