
Senior Data Engineer
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in United States.
• Design and implement high-volume data pipelines between Snowflake and Databricks utilizing PySpark and dbt.
• Manage pipeline scheduling, dependencies, monitoring, and automated failure recovery with Apache Airflow.
• Oversee Databricks platforms, including IAM roles, Amazon S3 access, secrets, governance policies, and cluster configurations.
• Create and deploy vector-based retrieval architecture and AI-driven workflows for marketing decision automation.
• Operationalize MLOps with MLflow and monitor Lakehouse environments.
• Develop NER, predictive models, deep learning architectures, and time-series forecasting models.
• Handle CI/CD pipelines using Git and Tekton.
• Construct and publish container images with buildah and skopeo.
• Deploy and manage containerized models and applications on Red Hat OpenShift/Kubernetes.
• Administer network routes, TLS termination, and high-availability services.
• Lead enterprise security compliance assessments across 20+ controls.
• Conduct SAST, vulnerability scanning, Privacy Impact Assessments, and STRIDE-based threat modeling.
• Work collaboratively with information security teams to address vulnerabilities, support compliance audits, and maintain Splunk logging and monitoring.
• Bachelor's degree in Computer Science or a related discipline with five (5) years of experience, or a Master's degree with three (3) years of experience.
• U.S. or foreign equivalent degree.
• A minimum of three (3) years of experience in architecting high-volume Snowflake-to-Databricks data pipelines using PySpark and dbt.
• At least three (3) years of experience in orchestrating workflows with Apache Airflow, including DAG dependencies, failure recovery, and monitoring.
• Minimum of three (3) years of experience administering Databricks, Amazon S3 IAM access, platform vaults, OpenShift secrets, governance, and cluster policies.
• A minimum of three (3) years of experience in building NER systems with TextBlob, gensim, and fastText.
• At least three (3) years of experience in developing models using XGBoost and Scikit-learn.
• Minimum of three (3) years of experience in constructing Keras deep learning architectures and time-series forecasting models.
• At least three (3) years of experience implementing information retrieval systems utilizing Transformer architectures and Transfer Learning.
• A minimum of three (3) years of experience with MLflow, automated model monitoring, Git, Tekton, buildah, skopeo, and Red Hat OpenShift/Kubernetes.
• At least three (3) years of experience with enterprise security compliance assessments, SonarQube, Qualys, Privacy Impact Assessments, and STRIDE threat modeling.
• Remote telecommuting options available.
• Eligibility for bonuses, commissions, and/or equity.
• Flexible work environments.
• Reasonable accommodations provided for job applicants.
Progress Partners
FYUL
CarringtonCrisp
Get handpicked remote jobs straight to your inbox weekly.