
IT Infrastructure Engineer – AI
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United Kingdom.
• Establish, configure, and oversee MCP (Model Context Protocol) servers that link AI systems with internal tools and data sources, which includes defining access, managing credentials, and resolving integration issues.
• Create, provision, and maintain hosting environments for in-house developed AI applications, ranging from lightweight PaaS deployments to comprehensive cloud infrastructure on AWS/GCP (VPCs, compute, containers, load balancing) based on the complexity and scale needs of the application.
• Have an understanding of AI functionalities within existing SaaS platforms (e.g., Atlassian Rovo, Slack AI, Okta AI, ChatGPT, Google Gemini).
• Provide and support infrastructure for AI agents/applications: managing service accounts, API keys, permissions, monitoring rate limits, and auditing access for agents operating across internal systems and cloud environments.
• Develop and sustain CI/CD pipelines and infrastructure-as-code (e.g., Terraform, CloudFormation) to ensure secure and repeatable deployment of AI-hosted applications and their supporting services.
• Keep documentation updated for all AI integrations and hosted environments, which includes architecture diagrams, access maps, runbooks, cost/usage tracking, and troubleshooting guides.
• Remain updated on AI tools and the cloud ecosystem, while proactively identifying beneficial tools, platforms, or integrations for the IT team or the wider organization.
• Assist in the management of UserTesting’s IT cloud infrastructure (AWS and/or GCP) and DevOps toolset throughout the organization concerning IT’s AI Infrastructure.
• Become a subject matter expert in AI-hosted platforms, cloud integrations, and supporting infrastructure.
• Monitor the uptime, performance, and costs of hosted AI applications; set up alerting and observability (e.g., Datadog, CloudWatch, GCO) to preemptively address issues before they impact users.
• Identify, prioritize, and mitigate technical debt within cloud infrastructure, deployment pipelines, and integrations.
• Collaborate with other IT System Administrators and Engineering/Security teams to enhance the reliability, automation, documentation, and supportability of hosted environments.
• Serve as an escalation point for intricate infrastructure and hosting challenges for the IT support team.
• Work with cross-functional teams (Security, Engineering, People Ops, etc.) to enhance cloud governance, security posture, and the adoption of AI tools/apps.
• Manage stakeholders; this role will occasionally function as a project lead. The ability to engage, demonstrate, and collaborate with non-technical counterparts is essential.
• Document procedures, system modifications, and troubleshooting workflows.
• Over 5 years of experience in cloud infrastructure and/or DevOps engineering, ideally with direct experience in AWS and/or GCP (compute, networking/VPCs, IAM, storage).
• Practical experience in deploying and hosting applications across cloud infrastructure (AWS/GCP).
• Hands-on experience with AI concepts such as MCP, AI Agents, and other AI-related technologies.
• Strong scripting and automation skills in Python and/or Bash, working proficiency with Git for version control, and familiarity with JSON/YAML for configuration and API tasks.
• Experience with containerization (Docker/Podman) and container orchestration (e.g., Kubernetes, ECS, or Cloud Run), including running and troubleshooting containerized services; experience with MCP servers or similar tools is a plus.
• Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Pulumi) and CI/CD pipelines (e.g., GitHub Actions).
• Familiarity with API-based integrations, including managing keys, scoping permissions, reading documentation, and troubleshooting connectivity issues.
• Ability to train and mentor IT staff on infrastructure and DevOps practices as they relate to your role and function.
• Understanding of IDP tools like Okta, along with the concepts of OAuth and SSO from an infrastructure/security standpoint.
• Certifications in cloud or AI platforms are beneficial (e.g., AWS, Google Cloud, Anthropic), but demonstrated practical experience is more highly valued; be ready to discuss this experience with real-world examples.
• Capacity to adapt in a constantly evolving environment.
• Competitive salary
• Flexible working hours
• Professional development budget
• Home office setup allowance
• Global team events
Learning Technologies Group plc
University of Wisconsin-Madison
NVIDIA
Astreya
Get handpicked remote jobs straight to your inbox weekly.