
Data Engineer, ML
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Construct the data and model layer for an AI-driven decision-support system within an accredited setting.
• Design and oversee the creation of graphs that represent entities, records, and their typed relationships.
• Execute probabilistic matching, blocking, candidate generation, pairwise scoring, clustering, threshold policies, deduplication, and known-record suppression.
• Ensure provenance so that every node and edge can be traced back to its asserting source.
• Create models for relevance and priority, calibration, threshold design, abstention policies, feature engineering, baselines, and error analysis.
• Implement embeddings, vector storage, retrieval, managed language-model integrations, and options for self-hosted or open-weight alternatives, prompts, and output schemas.
• Link generated text to cited source records while managing model packaging, serving, versioning, and rollback.
• Develop secure methods for ingestion, transformation, validation, and publishing of structured, semi-structured, and unstructured data.
• Perform quality checks, schema validation, lineage capture, audit logging, and source-drift detection.
• Produce statistically representative synthetic data for development prior to accessing live data.
• Collaborate with the Data Lead to adhere to data models and standards while aiding customer reviews.
• Document assumptions, caveats, transformation logic, and known limitations.
• Implement telemetry, maintain audit trails, and submit changes through a controlled release process.
• Provide explainable matching decisions, characterized model miss rates, grounded explanations, traceable pipelines, and consistent development progress.
• Over 5 years of experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering.
• Proficient in Python and SQL, especially with large and imperfect operational data.
• Experience in delivering systems intended for sustained operational use rather than just exploratory analysis.
• Regular use of AI-assisted development, demonstrating sound judgment regarding verification.
• Capability to articulate technical decisions to stakeholders who are accountable for them.
• US Citizenship is required.
• Active US Secret clearance is necessary.
• Favorable investigation and CAC eligibility from the outset.
• Elevated personnel security requirements are applicable to certain aspects of the work.
• Willingness to travel up to 25% to customer locations, DEFCON AI HQ, and vendor facilities.
• Preferred: active Top Secret clearance.
• Preferred experience includes probabilistic matching, record linkage, master data management, identity management, graph data modeling, PostgreSQL, pgvector, graph algorithms, model calibration, threshold design, cost-sensitive learning, scikit-learn, XGBoost, PyTorch, retrieval-augmented generation, prompt and output-schema design, self-hosted or open-weight models, fine-tuning, adapters, custom embeddings, AWS Glue, Airflow, dbt, Spark, Kafka, NiFi, document ingestion, synthetic test data, federal DevSecOps, RMF, ATO, DoW cloud environments, hardened base images, accreditation and deployment, sensitive federal or defense data, model cards, fairness testing, model monitoring, and NIST AI RMF.
• A fully remote, results-oriented work environment.
• Competitive salary, bonus, and equity package.
• Comprehensive health insurance, including medical, dental, and vision coverage for you and your family, fully paid by the employer.
• Unlimited PTO, subject to manager’s approval.
• A flexible work environment that allows you to manage your workday.
• 14 weeks of fully-paid parental leave.
Pluribus Digital
GoMining
Get handpicked remote jobs straight to your inbox weekly.