
Data Engineer – Onboarding
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in United States, +1 more state.
• Take ownership of the data ingestion layer that integrates device telemetry, transaction events, KYC/identity signals, and third-party enrichment into the platform by designing streaming pipelines (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc) that are reliable, observable, and cost-effective to scale.
• Develop and enhance our feature platform, where consistent Chronon feature definitions are processed by Flink for streaming and Spark for batch, utilizing aggregation windows ranging from one hour to 300 days, delivered to the rules engine and models within a sub-second time frame.
• Establish feature correctness as a core engineering discipline, including streaming vs. batch reconciliation, recomputation testing against the data warehouse, train/serve parity checks, and drift monitoring to identify broken features before they impact analysts.
• Productionize fraud and identity machine learning models through training pipelines on Vertex AI and Kubeflow, employing gradient-boosted and tree-based models (XGBoost, LightGBM, CatBoost, scikit-learn), hyperparameter optimization, SHAP-based explanations, and score normalization, while creating the automated retraining, champion/challenger promotion, and rollback systems we currently lack.
• Engineer KYC, AML, and identity risk signals, including document verification and doc-KYC outcomes, sanctions/PEP/adverse-media screening results, email and phone risk indicators, synthetic identity signals, bank and account verification, and periodic customer due diligence, transforming noisy, multi-vendor, multi-jurisdiction data into model-learnable features.
• Integrate and fortify new data sources, enlisting over 30 third-party enrichment providers operating concurrently on the request path, alongside our cross-client consortium network — managing failover behavior, timeout budgets, graceful degradation, caching, and cost.
• Oversee the warehouse and modeling layer in BigQuery, including partitioning strategy, staging-to-mart layer architecture, training datasets, and the ongoing migration from dbt to scheduled SQL and Python pipelines.
• Design the entity resolution and graph data structures that connect customers, devices, emails, phones, cards, bank accounts, and crypto addresses across clients, involving large-scale connected-components work.
• Ensure platform safety by design: implement field-level encryption for sensitive identifiers, enforce regional data residency in pipeline definitions, establish PII handling and deletion paths, and create feature-level gating to disable harmful signals without redeployment.
• Set the technical vision and elevate the team's capabilities — create design documentation, conduct reviews, mentor engineers and data scientists, and determine whether to build or buy solutions.
• A minimum of 8 years of experience in building production data and machine learning systems, demonstrating genuine ownership of both pipeline and model components. You have successfully deployed models that make significant automated decisions, not merely dashboards.
• Proficiency in Python and strong SQL skills. You are experienced with a distributed processing framework (Spark, Beam, or Flink) and can effectively reason about streaming semantics — such as windowing, watermarks, late data, exactly-once vs. at-least-once, and points of failure in correctness.
• Practical experience with a modern cloud data stack, preferably GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or equivalent AWS services, along with Docker, Kubernetes, Terraform, and CI/CD practices.
• Significant depth in practical machine learning engineering: feature stores and pipelines, training vs. serving skew, gradient-boosted tree models, handling class imbalance and rare-event modeling, tuning for thresholds and costs, model monitoring and drift detection, as well as explainability.
• Familiarity with high-volume, low-latency serving environments where feature fetching takes mere hundreds of milliseconds without a retry budget.
• Domain knowledge in fraud, risk, payments, lending, or identity/KYC, or the proven ability to quickly become proficient in a regulated domain. You understand the unique challenges of fraud modeling, including label latency, feedback loops, and adversarial drift compared to standard supervised learning.
• Comfort navigating data governance in a regulated context: PII, encryption, access control, regional data residency, and auditability.
• Excellent written communication skills. You can articulate a modeling trade-off to a fraud analyst and explain a pipeline design to a backend engineer, ensuring thorough documentation of your work.
• A proactive attitude and the ability to thrive in ambiguous situations. Much of this role involves determining what should be created and then executing on that vision.
• Competitive compensation package including cash and equity.
• Early exercise options available for all options, including pre-vested options.
• Work from any location: Emphasis on a remote-first culture.
• Flexible paid time off and a year-end break.
• Health insurance, dental, and vision coverage provided for employees and their dependents - *specific to the US and Canada*.
• 4% matching contributions in 401k / RRSP plans - *specific to the US and Canada*.
• A MacBook Pro will be shipped directly to you.
• One-time stipend for setting up your home office — desk, chair, screen, etc.
• Monthly meal stipend.
• Monthly stipend for social meet-ups.
• Annual health and wellness stipend.
• Annual learning stipend.
Snowflake
Cisco
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.