
Senior Software Engineer β Semantic Data Lake
Posted Aug 12

Posted Aug 12
This is a fully remote position, open to applicants in India.
β’ Architect and deploy the essential infrastructure for the Semantic Data Control Plane.
β’ Manage the data lifecycle encompassing ingestion, transformation, and final distribution.
β’ Develop observability frameworks and automated quality gates to ensure proactive platform reliability.
β’ Create self-service portals and standardized templates for domain teams.
β’ Automate governance and security protocols, including data lineage, RBAC, and PII protection.
β’ Establish foundational data architecture for AI-native systems, incorporating context and RAG-based systems.
β’ Utilize AI coding assistants to formulate transformation logic, generate tests, refactor pipelines, explore datasets, and create semantic documentation.
β’ Disseminate AI tooling patterns, prompts, and workflows within the Semantic Data Team.
β’ Partner with domain experts, data scientists, and product stakeholders to convert business concepts into frameworks.
β’ Execute classifications, KPIs, scoring algorithms, and business rules with traceability and data lineage.
β’ Contribute to data modeling, documentation, governance, and the responsible use of AI-generated code and artifacts.
β’ Collaborate across teams to integrate ingestion, MDM, and data product layers.
β’ 4β8 years of experience in data engineering or software engineering, focusing on data transformation, modeling, or analytics platforms.
β’ Strong expertise in SQL.
β’ Proficient in at least one general-purpose programming language, such as Python or Scala.
β’ Regular use of AI coding tools, including Claude, GitHub Copilot, Cursor, or similar technologies.
β’ Practical experience with LLM-driven architectures.
β’ Background in designing RAG systems.
β’ Experience in constructing agentic workflows like LangGraph or CrewAI.
β’ Familiarity with vector or graph databases for context and ontology.
β’ Knowledge of prompt design, Spec-Driven Development, and AI-assisted code review.
β’ Deep understanding of complex orchestration systems, particularly Airflow.
β’ Experience with data quality frameworks, such as Great Expectations.
β’ Proficient with Terraform, Kubernetes, and CI/CD automation.
β’ Experience implementing enterprise governance using DataHub or similar platforms.
β’ Familiar with end-to-end lineage, cataloging, and RBAC policies.
β’ Experienced in creating observability frameworks, including Grafana and xMatters integration.
β’ Skilled in building self-service portals and developer tools.
β’ Knowledge of data quality practices, including validation, enrichment, schema enforcement, and business rule encoding.
β’ Ability to work collaboratively in a cross-functional environment.
β’ Demonstrated track record of traceability, reproducibility, and semantic clarity.
β’ Comprehensive and competitive benefits package.
β’ Offerings designed to foster personal and professional well-being.
β’ Equal opportunity employer.
β’ Reasonable accommodations available for qualified individuals with disabilities.
β’ Drug-free workplace.
Atomic - Remote Jobs
Affidea
GoFasti
StructureIt
Get handpicked remote jobs straight to your inbox weekly.