
Senior Software Engineer – Semantic Data Lake
Posted Jun 18

Posted Jun 18
This is a fully remote position, open to applicants in California, +5 more states.
• Develop and implement semantically consistent, scalable 360 data models that unify data across various domains.
• Create and maintain transformation pipelines that apply cleansing, standardization, enrichment, and derived logic to datasets within domains.
• Produce production-quality, testable code in SQL and Python (or a similar language)—ensuring the delivery of efficient and maintainable data assets.
• Utilize AI coding assistants (such as Claude, Copilot, Cursor, and others) to expedite development—composing transformation logic, generating tests, refactoring pipelines, examining datasets, and creating semantic documentation—while critically assessing AI outputs for accuracy, performance, and alignment with business objectives.
• Collaborate closely with the data products team to comprehend business requirements and ensure that semantic models meet their specifications.
• Implement logic for classifications, KPIs, scoring algorithms, and business rules, ensuring traceability and data lineage.
• Work across teams to integrate with ingestion, MDM, and data product layers, and investigate opportunities to expose 360 objects to LLM-powered and agentic applications.
• 4-8 years of experience in data engineering or software engineering, emphasizing data transformation, modeling, or analytics platforms.
• Strong expertise in SQL and at least one general-purpose programming language such as Python or Scala.
• Proven experience as an AI-native engineer—utilizing tools like Claude, GitHub Copilot, Cursor, or similar as part of your daily development process.
• Familiarity with modern AI engineering practices such as prompt design, context engineering, AI-assisted code reviews, and integrating LLMs or AI agents into engineering or data workflows.
• Experience in building and scaling wide, entity-based tables and modeling domain concepts (like customer, fleet, provider) into robust data objects.
• Strong understanding of data quality practices—including validation, enrichment, schema enforcement, and encoding business rules.
• Experience handling large-scale datasets and optimizing transformation pipelines for performance and maintainability.
• Ability to operate in a collaborative, cross-functional environment, balancing business logic with platform scalability.
• A focus on traceability, reproducibility, and semantic clarity—creating data models that can be trusted and reused by both humans and AI systems.
• Bachelor's degree in Computer Science, Software Engineering, or a related field.
• Health, dental, and vision insurance.
• Retirement savings plan.
• Paid time off.
• Health savings account.
• Flexible spending accounts.
• Life insurance.
• Disability insurance.
• Tuition reimbursement.
• Commission under an applicable plan.
• Quarterly or annual bonuses based on role and applicable plan.
Phase2
Job Mobz
RTX
Get handpicked remote jobs straight to your inbox weekly.