
Application Reliability Engineer – GCP
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Provide production support and maintenance for enterprise applications hosted on Google App Engine.
• Triage, diagnose, restore service, and drive user-impacting incidents to resolution within SLAs/SLOs.
• Troubleshoot application errors, failed requests, latency issues, performance degradation, service failures, configuration problems, quota/scaling limits, and integration failures.
• Design, develop, and implement feature enhancements to existing applications.
• Support microservices architecture, APIs/contracts, authentication, inter-service communication, and failure/retry behavior.
• Oversee build, release, and deployment activities across development, test, pre-production, and production environments.
• Manage App Engine deployments, including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, configuration, and scaling.
• Conduct root-cause analysis and implement sustainable solutions.
• Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.
• Support platform, framework, library, dependency, and runtime upgrades.
• Manage IAM, service accounts, access controls, secrets management, and operational governance.
• Participate in change management, release-readiness reviews, and on-call/rotational support.
• Collaborate with clients and cross-functional teams in Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform.
• Adapt to client-specific engineering, security, review, change-management, and operational processes.
• Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.
• 3-7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related field.
• Strong hands-on experience in supporting, maintaining, and enhancing production applications on Google Cloud Platform.
• Practical experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting.
• Solid understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.
• Experience with multi-environment deployments and promotions, release validation, rollback, and change control.
• Proficient programming skills in one or more of the following: Python, Java, Node.js/JavaScript, or Go.
• Working knowledge of SQL and application data stores like Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.
• Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.
• Experience with CI/CD pipelines and automated build/deployment tools such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.
• Experience troubleshooting complex production environments and performing root-cause analysis under pressure.
• Ability to quickly comprehend existing systems, codebases, services, configurations, and client-specific tools and workflows.
• Strong communication and collaboration skills, especially during user-impacting incidents.
• Preferred: experience with large-scale enterprise applications, Cloud Run, GKE, Cloud Functions, Apigee/API Gateway, Pub/Sub, Cloud Tasks, Cloud Scheduler, Terraform, SRE practices, Docker/Kubernetes, frontend/full-stack applications, BI platforms, or AI/ML/GenAI/LLM applications including Vertex AI.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive company culture.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.