
Senior Data Engineer
Posted 1 hour ago

Posted 1 hour ago
This is a fully remote position, open to applicants in California.
• Design and sustain comprehensive data pipelines and backend ingestion processes, contributing to the development of Samsara's Data Platform for enhanced automation and analytics.
• Collaborate with various data sources such as ERP (Netsuite), CRM (Salesforce), Product, Order Flow, and Support ticket data.
• Oversee essential data pipelines to support growth initiatives and sophisticated analytics.
• Enable data integration and transformation for transferring data across applications, ensuring compatibility with data layers and the data lake.
• Enhance data architecture, data quality, monitoring, observability, and data accessibility.
• Create data transformations using SQL/Python to produce data products utilized by the Analytics, Marketing Operations, and Sales Operations teams.
• Architect, construct, and manage large-scale Spark and PySpark workflows for batch and streaming data processing on Databricks and cloud platforms.
• Improve Spark job performance through optimization techniques such as tuning partitioning, shuffle, caching, and resource allocation for reliable and efficient production operations.
• Develop, construct, and maintain data APIs in Python using frameworks like FastAPI.
• Manage API runtime within AWS ecosystems, including Lambda and RDS.
• Monitor and enhance API performance, tracking observability with tools like Data Dog or Splunk.
• Promote and exemplify Samsara's cultural principles as the company scales globally.
• Mentor junior team members, providing technical guidance, training, and knowledge-sharing across teams.
• Bachelor's degree in computer science, data engineering, data science, information technology, or a comparable engineering discipline.
• Over 10 years of professional experience as a Software Engineer with a focus on data, or as a Data Engineer.
• More than 8 years of experience in constructing and maintaining large-scale, production-grade end-to-end data pipelines, including Data Modeling.
• At least 5 years of practical experience with Spark/PySpark in a production setting, encompassing job optimization and performance tuning.
• 3 or more years of direct experience in developing and managing APIs, preferably in Python.
• Proficient programming skills in Python and SQL, alongside experience with cloud data warehouse/lakehouse technologies (e.g., Snowflake, Google BigQuery, Databricks, or Apache Iceberg).
• Familiarity with ETL tools such as Fivetran, DBT, or similar alternatives.
• Experience with Python-based API frameworks and API management tools.
• Knowledge of RDBMS technologies like MySQL, AWS RDS/Aurora, PostgreSQL, Oracle, MS SQL Server, or their equivalents.
• Proficient in cloud platforms: AWS, Azure, and/or GCP.
• Flexible, employee-driven remote work model
• Professional development stipend
• Comprehensive health and parental leave plans
Redhorse Corporation
GE Aerospace
Sayari
Minor Hotels Europe and Americas
Get handpicked remote jobs straight to your inbox weekly.