Remotery

Application Reliability Engineer – GCP

atInnoDataRemoteUS flagUnited StatesFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$60k – $110k/year

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Provide production support and maintenance for enterprise applications hosted on Google App Engine.

• Triage, diagnose, restore service, and drive user-impacting incidents to resolution within SLAs/SLOs.

• Troubleshoot application errors, failed requests, latency issues, performance degradation, service failures, configuration problems, quota/scaling limits, and integration failures.

• Design, develop, and implement feature enhancements to existing applications.

• Support microservices architecture, APIs/contracts, authentication, inter-service communication, and failure/retry behavior.

• Oversee build, release, and deployment activities across development, test, pre-production, and production environments.

• Manage App Engine deployments, including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, configuration, and scaling.

• Conduct root-cause analysis and implement sustainable solutions.

• Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.

• Support platform, framework, library, dependency, and runtime upgrades.

• Manage IAM, service accounts, access controls, secrets management, and operational governance.

• Participate in change management, release-readiness reviews, and on-call/rotational support.

• Collaborate with clients and cross-functional teams in Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform.

• Adapt to client-specific engineering, security, review, change-management, and operational processes.

• Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.


⛳️ Requirements

• 3-7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related field.

• Strong hands-on experience in supporting, maintaining, and enhancing production applications on Google Cloud Platform.

• Practical experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting.

• Solid understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.

• Experience with multi-environment deployments and promotions, release validation, rollback, and change control.

• Proficient programming skills in one or more of the following: Python, Java, Node.js/JavaScript, or Go.

• Working knowledge of SQL and application data stores like Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.

• Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.

• Experience with CI/CD pipelines and automated build/deployment tools such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.

• Experience troubleshooting complex production environments and performing root-cause analysis under pressure.

• Ability to quickly comprehend existing systems, codebases, services, configurations, and client-specific tools and workflows.

• Strong communication and collaboration skills, especially during user-impacting incidents.

• Preferred: experience with large-scale enterprise applications, Cloud Run, GKE, Cloud Functions, Apigee/API Gateway, Pub/Sub, Cloud Tasks, Cloud Scheduler, Terraform, SRE practices, Docker/Kubernetes, frontend/full-stack applications, BI platforms, or AI/ML/GenAI/LLM applications including Vertex AI.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Opportunities for professional development and continuous learning.

• Collaborative and inclusive company culture.

People also viewed

TEKsystems16 hours ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems16 hours ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data20 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data20 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data20 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data20 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers