
Senior Platform Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Develop, manage, support, and enhance enterprise observability platforms for both applications and infrastructure.
• Operate and fine-tune observability backends, ingestion pipelines, agents/collectors, and data lifecycle management processes.
• Execute observability solutions and standards established by Staff and Principal Engineers.
• Construct and maintain telemetry pipelines for logs, metrics, traces, and events.
• Lead the onboarding of telemetry and data for applications, infrastructure, Kubernetes, cloud, and platform teams.
• Offer product support to internal teams by addressing telemetry ingestion, query performance, data quality, and platform usage challenges.
• Monitor and enhance platform availability, performance, capacity, scalability, and reliability.
• Resolve complex production issues related to telemetry ingestion, processing, storage, indexing, and query performance.
• Optimize data stores, retention policies, indexing strategies, and storage utilization.
• Develop and maintain dashboards, alerts, integrations, and operational tools.
• Automate platform provisioning, configuration, and deployment using Infrastructure as Code and CI/CD practices.
• Collaborate with Application, SRE, Infrastructure, and Security teams to enhance telemetry quality, coverage, and incident response.
• Over 7 years of practical experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or Observability Engineering.
• Extensive hands-on experience with at least one major observability stack, such as Elastic, Grafana, Datadog, Splunk, or similar.
• Proven experience in operating and troubleshooting production observability platforms at scale.
• Strong comprehension of logs, metrics, distributed tracing, telemetry pipelines, and fundamental observability concepts.
• Experience in onboarding telemetry/data sources and providing support to internal customers through integration and troubleshooting.
• Practical experience with Kubernetes, containers, Linux, networking, and cloud infrastructure.
• Familiarity with OpenTelemetry or comparable instrumentation and collection frameworks.
• Advanced troubleshooting capabilities across applications, infrastructure, and distributed systems.
• Practical experience with Infrastructure as Code tools like Terraform.
• Experience in constructing CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, or similar tools.
• Proficient in automation and scripting using Python, Go, Bash, or similar languages.
• Experience in performance tuning, capacity planning, data lifecycle management, and platform optimization.
• Working knowledge of telemetry collection, signal processing, cross-signal analysis, alerting, and production troubleshooting.
• Ability to build, operate, troubleshoot, and enhance production-grade platform solutions.
• Experience in implementing SLIs, SLOs, alerting standards, and reliability monitoring.
• Familiarity with modern observability tools such as Grafana, Prometheus, ClickHouse, or pipeline/stream processing platforms.
• Experience with log/metric/trace collectors and streaming technologies like Fluent Bit and Kafka.
• Background in high-volume telemetry environments and cost optimization strategies.
• Experience with GitOps workflows and observability-as-code methodologies.
• Ability to create onboarding documentation, runbooks, and self-service guidance for platform users.
• Equal Opportunity Employer.
• Inclusive workforce and diversity-focused work environment.
Grafana Labs
Grafana Labs
Grafana Labs
Get handpicked remote jobs straight to your inbox weekly.