
Senior Platform Engineer
Posted Aug 13

Posted Aug 13
This is a fully remote position, open to applicants in India.
β’ Develop, manage, support, and enhance enterprise observability platforms for both applications and infrastructure.
β’ Operate and fine-tune observability backends, ingestion pipelines, agents/collectors, and data lifecycle management processes.
β’ Execute observability solutions and standards established by Staff and Principal Engineers.
β’ Construct and maintain telemetry pipelines for logs, metrics, traces, and events.
β’ Lead the onboarding of telemetry and data for applications, infrastructure, Kubernetes, cloud, and platform teams.
β’ Offer product support to internal teams by addressing telemetry ingestion, query performance, data quality, and platform usage challenges.
β’ Monitor and enhance platform availability, performance, capacity, scalability, and reliability.
β’ Resolve complex production issues related to telemetry ingestion, processing, storage, indexing, and query performance.
β’ Optimize data stores, retention policies, indexing strategies, and storage utilization.
β’ Develop and maintain dashboards, alerts, integrations, and operational tools.
β’ Automate platform provisioning, configuration, and deployment using Infrastructure as Code and CI/CD practices.
β’ Collaborate with Application, SRE, Infrastructure, and Security teams to enhance telemetry quality, coverage, and incident response.
β’ Over 7 years of practical experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or Observability Engineering.
β’ Extensive hands-on experience with at least one major observability stack, such as Elastic, Grafana, Datadog, Splunk, or similar.
β’ Proven experience in operating and troubleshooting production observability platforms at scale.
β’ Strong comprehension of logs, metrics, distributed tracing, telemetry pipelines, and fundamental observability concepts.
β’ Experience in onboarding telemetry/data sources and providing support to internal customers through integration and troubleshooting.
β’ Practical experience with Kubernetes, containers, Linux, networking, and cloud infrastructure.
β’ Familiarity with OpenTelemetry or comparable instrumentation and collection frameworks.
β’ Advanced troubleshooting capabilities across applications, infrastructure, and distributed systems.
β’ Practical experience with Infrastructure as Code tools like Terraform.
β’ Experience in constructing CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, or similar tools.
β’ Proficient in automation and scripting using Python, Go, Bash, or similar languages.
β’ Experience in performance tuning, capacity planning, data lifecycle management, and platform optimization.
β’ Working knowledge of telemetry collection, signal processing, cross-signal analysis, alerting, and production troubleshooting.
β’ Ability to build, operate, troubleshoot, and enhance production-grade platform solutions.
β’ Experience in implementing SLIs, SLOs, alerting standards, and reliability monitoring.
β’ Familiarity with modern observability tools such as Grafana, Prometheus, ClickHouse, or pipeline/stream processing platforms.
β’ Experience with log/metric/trace collectors and streaming technologies like Fluent Bit and Kafka.
β’ Background in high-volume telemetry environments and cost optimization strategies.
β’ Experience with GitOps workflows and observability-as-code methodologies.
β’ Ability to create onboarding documentation, runbooks, and self-service guidance for platform users.
β’ Equal Opportunity Employer.
β’ Inclusive workforce and diversity-focused work environment.
Welyk
Trase
hims & hers
Hummingbird Healthcare
Get handpicked remote jobs straight to your inbox weekly.