
Systems Administrator
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Canada.
• Oversee GCP/AWS cloud initiatives, including IAM, service accounts, networking, storage, compute, quotas, and SaaS environments.
• Manage Cloud Run, Docker containers, registries, DNS, TLS, load balancing, and Google Cloud IAP operations.
• Administer identities, RBAC, SSO, OAuth, service accounts, machine-to-machine authentication, secrets, API keys, tokens, and certificates.
• Uphold approved AI agents, MCP servers, tools, integrations, data sources, permissions, and credentials.
• Oversee Dockerfiles, container images, Python/application dependencies, registries, runtime configurations, and rollback procedures.
• Track logs, metrics, traces, alerts, dashboards, health checks, API limits, model/tool failures, token usage, cloud expenses, and service uptime.
• Maintain CI/CD pipelines and infrastructure-as-code, encompassing source control, security scanning, testing, approvals, and deployment traceability.
• Conduct incident triage, escalation, root-cause analysis, backup/recovery, disaster recovery, and service continuity measures.
• Manage incidents, service requests, problems, and changes utilizing ServiceNow and Jira.
• Assist in vulnerability remediation, patching, security investigations, audits, threat modeling, and risk assessments.
• Identify shadow AI services, unmanaged integrations, unapproved MCP servers, and overprivileged identities.
• Develop operational runbooks, technical documentation, and lifecycle/ownership documentation.
• Facilitate the transition of scientist-managed prototypes into secure, maintainable production deployments.
• Support the secure, reliable, and compliant operation of cloud-hosted research applications, agentic AI systems, scientific SaaS platforms, and production services.
• Collaborate across GCP, AWS, cloud containers, AI/ML models, MCP servers, APIs, scientific databases, and third-party SaaS platforms.
• Assist scientists, developers, IT, cybersecurity, networking, data teams, architecture, and vendors in transforming research prototypes into secure, repeatable, production-ready services.
• Bachelor’s degree in Computer Science, Information Systems, Engineering, Technology, or a related field.
• Over 3 years of experience in systems/cloud administration.
• Experience in production administration with GCP, AWS, Azure, or similar cloud platforms.
• Proficient with Linux.
• Practical knowledge of Docker/containers.
• Experience with Python environments.
• Familiarity with APIs.
• Knowledge of networking, DNS, and TLS.
• Experience with IAM, SSO, OAuth, and service accounts.
• Hands-on experience with secrets management.
• Experience in serverless/container orchestration such as Cloud Run or Kubernetes.
• Understanding of CI/CD methodologies.
• Knowledge of Infrastructure as Code.
• Familiarity with monitoring and observability practices.
• Understanding of incident management principles.
• Knowledge of backup/recovery strategies.
• Experience with patching processes.
• Familiarity with vulnerability remediation techniques.
• Experience with ServiceNow and Jira.
• Strong skills in troubleshooting, documentation, communication, and cross-functional collaboration.
• Remote work arrangement.
Sprinter Health
Ventra Health
Midnite
Get handpicked remote jobs straight to your inbox weekly.