
Senior Software Engineer – Semantic Data Lake
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Architect and deploy the essential infrastructure for the Semantic Data Control Plane.
• Manage the data lifecycle encompassing ingestion, transformation, and final distribution.
• Develop observability frameworks and automated quality gates to ensure proactive platform reliability.
• Create self-service portals and standardized templates for domain teams.
• Automate governance and security protocols, including data lineage, RBAC, and PII protection.
• Establish foundational data architecture for AI-native systems, incorporating context and RAG-based systems.
• Utilize AI coding assistants to formulate transformation logic, generate tests, refactor pipelines, explore datasets, and create semantic documentation.
• Disseminate AI tooling patterns, prompts, and workflows within the Semantic Data Team.
• Partner with domain experts, data scientists, and product stakeholders to convert business concepts into frameworks.
• Execute classifications, KPIs, scoring algorithms, and business rules with traceability and data lineage.
• Contribute to data modeling, documentation, governance, and the responsible use of AI-generated code and artifacts.
• Collaborate across teams to integrate ingestion, MDM, and data product layers.
• 4–8 years of experience in data engineering or software engineering, focusing on data transformation, modeling, or analytics platforms.
• Strong expertise in SQL.
• Proficient in at least one general-purpose programming language, such as Python or Scala.
• Regular use of AI coding tools, including Claude, GitHub Copilot, Cursor, or similar technologies.
• Practical experience with LLM-driven architectures.
• Background in designing RAG systems.
• Experience in constructing agentic workflows like LangGraph or CrewAI.
• Familiarity with vector or graph databases for context and ontology.
• Knowledge of prompt design, Spec-Driven Development, and AI-assisted code review.
• Deep understanding of complex orchestration systems, particularly Airflow.
• Experience with data quality frameworks, such as Great Expectations.
• Proficient with Terraform, Kubernetes, and CI/CD automation.
• Experience implementing enterprise governance using DataHub or similar platforms.
• Familiar with end-to-end lineage, cataloging, and RBAC policies.
• Experienced in creating observability frameworks, including Grafana and xMatters integration.
• Skilled in building self-service portals and developer tools.
• Knowledge of data quality practices, including validation, enrichment, schema enforcement, and business rule encoding.
• Ability to work collaboratively in a cross-functional environment.
• Demonstrated track record of traceability, reproducibility, and semantic clarity.
• Comprehensive and competitive benefits package.
• Offerings designed to foster personal and professional well-being.
• Equal opportunity employer.
• Reasonable accommodations available for qualified individuals with disabilities.
• Drug-free workplace.
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.