
Staff Software Engineer β Semantic Foundation
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in California, +5 more states.
β’ Design, implement, scale, and sustain highly available Apache Airflow clusters utilizing Helm, Kubernetes, and Terraform.
β’ Lead DevOps initiatives using Terraform, Helm Charts, and ArgoCD.
β’ Automate platform deployments, oversee self-hosted GitHub runners, and uphold GitOps workflows.
β’ Architect and manage CI/CD pipelines through GitHub Actions for data pipelines, infrastructure elements, and DAGs.
β’ Provision and manage Kubernetes EKS/AKS clusters, Docker containers, and AWS/Azure cloud infrastructure.
β’ Develop and maintain monitoring, logging, and alerting systems using Grafana, Prometheus, Loki, and centralized log management solutions.
β’ Construct scalable data infrastructure for Apache Spark, AWS EMR, Snowflake, Azure Synapse, Apache Kafka, and dbt.
β’ Deploy and maintain data discovery and metadata tools such as DataHub.
β’ Monitor, audit, and optimize AWS and Azure compute and storage expenditures.
β’ Create internal tools, CLI utilities, and dynamic workflow templates.
β’ Take ownership of data platform modules from architecture through deployment, production operations, and incident management.
β’ Establish and uphold engineering standards related to code quality, design, testing, data lineage, security, access controls, IAM, and schema management.
β’ Collaborate with data leads, product managers, and business stakeholders on a 1β2 year data platform roadmap.
β’ Mentor junior and mid-level data platform engineers.
β’ Over 6 years of practical experience in Data Platform Engineering, DevOps, or Site Reliability Engineering within large-scale data environments.
β’ Extensive production experience in managing, tuning, dynamically scaling, and troubleshooting Apache Airflow infrastructure.
β’ Expert proficiency in Kubernetes, Helm, Terraform, Docker, ArgoCD, and GitHub Actions, including custom runner setups.
β’ Demonstrated experience in configuring production alerting, metrics collection, and log aggregation using Grafana, Prometheus, and Loki.
β’ In-depth operational and configuration expertise with Spark, AWS EMR, Snowflake, Azure Synapse, dbt, and Apache Kafka.
β’ Hands-on experience with AWS and/or Azure, showcasing a history of optimizing cloud costs.
β’ Experience in deploying or managing data cataloging tools like DataHub, Amundsen, or similar metadata management platforms.
β’ Strong programming abilities in Python, Bash, Go, or SQL.
β’ Capable of transforming complex architectural requirements into production-grade deployments.
β’ Committed to automating manual operational tasks with code, automated tests, automated CI/CD checks, and resilient self-healing infrastructure.
β’ Must reside within 30 miles of Portland, ME; Boston, MA; Chicago, IL; Dallas, TX; San Francisco Bay Area, CA; or Seattle, WA.
β’ Health, dental, and vision insurance.
β’ Retirement savings plan.
β’ Paid time off.
β’ Health savings account.
β’ Flexible spending accounts.
β’ Life insurance.
β’ Disability insurance.
β’ Tuition reimbursement.
β’ Eligibility for quarterly or annual bonuses for non-sales roles.
Arista Networks
knowmad mood
Capital One
HealthEdge
Get handpicked remote jobs straight to your inbox weekly.