
Staff Engineer – Data Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Brazil.
• Create logical and physical data models for unstructured and semi-structured content derived from KM pipelines.
• Define the boundaries and ownership of data products within specific domains.
• Establish standards for metadata and tagging taxonomies across various knowledge resources.
• Assign and uphold security and sensitivity classifications in accordance with governance, privacy, and legal/risk standards.
• Register, document, and maintain data products in the Databricks Unity Catalog, including schemas, access permissions, lineage, and catalog-level metadata.
• Collaborate with data engineers to synchronize ingestion, transformation, and storage methods with the modeled domain structures.
• Work alongside stakeholders from Knowledge Products, Research Products, and Architecture/Data/Technology to address downstream consumption requirements.
• Assist in privacy and legal review processes through the classification and documentation of data products.
• Develop and document repeatable modeling standards and playbooks.
• Deliver domain models, metadata taxonomies, registered and discoverable data products, security classifications, and repeatable modeling standards within the initial 6–12 months.
• Minimum of 5 years of experience in data modeling, data architecture, or information architecture.
• Significant exposure to unstructured or semi-structured data.
• Direct experience in or related to Knowledge Management, content management, or enterprise search.
• Practical experience with a modern data catalog; familiarity with Databricks Unity Catalog is highly preferred.
• Capability to define data domains and product boundaries in a large, multi-stakeholder environment.
• Hands-on knowledge of metadata management, tagging schemas, taxonomies, controlled vocabularies, or ontology design.
• Understanding of data security and sensitivity classification frameworks along with access control in a lakehouse environment.
• Experience collaborating with data engineering teams on ingestion and pipeline design.
• Excellent written and verbal communication skills.
• Preferred experience with enterprise knowledge platforms or AI-driven retrieval systems.
• Familiarity with Databricks Delta Lake, Delta Sharing, or Lakehouse Federation is preferred.
• Previous experience in professional services, consulting, or document/case-intensive knowledge environments is preferred.
• Exposure to Legal/Risk/Privacy review processes is preferred.
• A background in library science, information science, or applied ontology is advantageous but not mandatory.
• Must possess strong Data Modeling and Databricks skills.
• Opportunity for remote work.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.