
Senior Data Science Engineer
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in United States.
β’ Take ownership of the complete lifecycle from the ingestion of raw data to the deployment of models and the evaluation of their real-world business impact.
β’ Conduct research, prototype, and develop machine learning and large language model-based solutions for intricate business challenges, with an emphasis on risk detection and prioritization.
β’ Package models into production-ready APIs and seamlessly integrate them into the primary SaaS product.
β’ Ensure that model outputs are interpretable by converting predictions into actionable reason codes.
β’ Collaborate with operational teams to collect feedback, enhance features, and boost model relevance.
β’ Design, construct, and sustain scalable pipelines that ingest data from various sources into the data warehouse/lake.
β’ Execute data validation, quality checks, and transformation workflows across raw, curated, and serving layers.
β’ Create and maintain curated datasets for analytics and model training purposes.
β’ Implement and uphold CI/CD pipelines for data workflows and the deployment of ML models.
β’ Monitor pipeline latency, data drift, and model performance; design alerting mechanisms and retraining triggers.
β’ Establish success metrics, monitor ROI, and iterate models based on their real-world effectiveness.
β’ Oversee infrastructure as code and containerized deployments for consistent releases.
β’ 5β8+ years of experience in data engineering and data science/machine learning, with a proven history of deploying models to production.
β’ Proficient in Python programming.
β’ Experience using Spark/PySpark for large-scale data processing tasks.
β’ Advanced SQL skills for complex transformations, analyses, and data modeling.
β’ Practical experience with cloud data platforms such as Databricks or Snowflake.
β’ Familiarity with ETL/ELT frameworks, including dbt, Lakeflow Declarative Pipelines, Databricks Autoloader, Informatica, or similar tools.
β’ Knowledge of ML experiment tracking tools such as MLflow or Weights & Biases.
β’ Fluency in DevOps practices: Git-based development, branching strategies, CI/CD, infrastructure as code (DABs/Terraform), and Docker.
β’ Experience with orchestration tools like Databricks Workflows or Apache Airflow.
β’ Strong plus: hands-on experience with large language models and Generative AI techniques in a production setting (prompt engineering, retrieval-augmented generation architectures, fine-tuning, or evaluation frameworks).
β’ Strong plus: experience in building or managing ML platforms, feature stores, or model registries.
β’ Strong plus: prior experience in risk, compliance, fraud detection, or other critical ML domains.
β’ Visa sponsorship is not available for this position.
β’ International remote work is not supported.
β’ Competitive compensation package.
β’ Flexible work options available.
β’ A team that is genuinely invested in your success.
Lightcast
Allstate
ACT
Working Families Party
Get handpicked remote jobs straight to your inbox weekly.