
Lead Application Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Canada.
• Provide production support and maintenance for enterprise applications hosted on Google App Engine.
• Triage, diagnose, restore service, and resolve user-impacting incidents within agreed SLAs/SLOs.
• Troubleshoot application errors, latency issues, performance degradation, service failures, configuration problems, quota/scaling limits, and integration failures.
• Design, develop, and implement enhancements to existing applications.
• Support microservices architecture, APIs, authentication, inter-service communication, and failure/retry behavior.
• Oversee build, release, and deployment activities across development, testing, pre-production, and production environments.
• Manage App Engine versions, traffic splitting/migration, canary and staged rollouts, rollbacks, configuration, and scaling.
• Conduct root-cause analysis and apply sustainable fixes.
• Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.
• Support platform, framework, library, dependency, and runtime upgrades.
• Handle IAM, service accounts, access controls, secrets management, and operational governance.
• Participate in change management, release-readiness reviews, and on-call/rotational support.
• Collaborate with clients and cross-functional teams including Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams.
• Create and maintain technical documentation, runbooks, troubleshooting guides, deployment procedures, and support playbooks.
• 3–7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related field.
• Strong hands-on experience in supporting production applications on Google Cloud Platform.
• Practical experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting.
• Understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.
• Experience with multi-environment deployments, release validation, rollback, and change control.
• Proficient programming skills in one or more of Python, Java, Node.js/JavaScript, or Go.
• Working knowledge of SQL and application data stores including Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.
• Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.
• Experience with CI/CD pipelines and automated build/deployment tools such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.
• Experience troubleshooting complex production environments and conducting root-cause analysis under pressure.
• Ability to comprehend existing systems, codebases, services, configurations, and client-specific workflows.
• Strong communication and collaboration skills.
• Preferred: experience with large-scale enterprise applications, Cloud Run, GKE, Cloud Functions, Apigee/API Gateway, Pub/Sub, Cloud Tasks, Terraform, SRE practices, Docker, Kubernetes, frontend/full-stack applications, BI platforms, or AI/ML applications on GCP.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional development and career growth.
• Dynamic work environment with a focus on innovation and collaboration.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.