
Data & ML Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Develop the data and model infrastructure behind an AI-supported decision-making system within a certified setting.
• Create and manage graphs that represent entities, records, and their typed relationships.
• Execute probabilistic matching, deduplication, known-record suppression, and source provenance.
• Construct relevance and priority models, calibration, thresholds, abstention policies, feature engineering, baselines, and error analysis.
• Implement embeddings, vector storage, retrieval, language model integrations, prompts, output schemas, grounded generation, model serving, versioning, and rollback.
• Establish secure ingestion, transformation, validation, and publishing pipelines across structured, semi-structured, and unstructured data sources.
• Conduct quality checks, schema validation, lineage capture, audit logging, and source-drift detection.
• Create statistically representative synthetic data for development prior to accessing live data.
• Document assumptions, caveats, transformation processes, known limitations, interfaces, and data flows.
• Instrument telemetry and maintain audit trails for recommendations, human interventions, and model versions.
• Submit modifications to models and pipelines through a controlled release process.
• Adhere to the standards established by the Data Lead, who approves designs and oversees them through customer review.
• 5+ years of experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering.
• Proficient in Python and SQL, with proven experience handling large, imperfect operational datasets.
• Experience in delivering systems for ongoing operational use, not just exploratory analysis.
• Regular use of AI-assisted development, with sound judgment on its value and the necessity for output verification.
• Capability to articulate technical decisions to stakeholders who need to defend them without in-depth understanding.
• US Citizenship Required.
• Active US Secret clearance.
• Favorable investigation and CAC eligibility from the outset.
• Elevated personnel security requirements are applicable to certain aspects of the role.
• Willingness to travel up to 25% to customer locations, DEFCON AI HQ, and vendor facilities.
• Preferred: active Top Secret clearance.
• Preferred experience with probabilistic matching, record linkage, master data management, identity management, graph data modeling, PostgreSQL, pgvector, graph algorithms, model calibration, threshold design, cost-sensitive learning, scikit-learn, XGBoost, PyTorch, retrieval-augmented generation, prompt and output-schema design, self-hosted or open-weight models, fine-tuning, adapters, custom embeddings, AWS Glue, Airflow, dbt, Spark, Kafka, NiFi, document ingestion, synthetic test data, federal DevSecOps, RMF, ATO, DoW cloud environments, hardened base images, accreditation and deployment, sensitive federal or defense data, model cards, fairness testing, model monitoring, and NIST AI RMF.
• Competitive salary, bonus, and equity package.
• 100% employer-covered, comprehensive health insurance including medical, dental, and vision for you and your family.
• Unlimited PTO, subject to manager’s approval.
• Flexible work environment allowing you to manage your workday.
• 14 weeks of fully-paid parental leave.
• Career progression opportunity with potential for rapid advancement based on strong performance as the firm expands.
• Opportunities for professional development and ongoing learning.
• Optional 401K, FSA, and equity incentives available.
• Mental health benefits accessible through Tara Mind.
• Cost-effective GLP-1 solutions available through Crux.
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.