
Senior Knowledge Graph Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Poland.
• Design, develop, and maintain high-speed GraphRAG ingestion pipelines that convert relational data (ERP, SQL), unstructured ESG reports, and streaming feeds into operational Labeled Property Graphs (Neo4j, Memgraph) and RDF Triple Stores.
• Create automated Named Entity Recognition (NER), entity linking, and deduplication workflows to streamline mismatched vendor profiles, material SKUs, and facility coordinates into unified canonical graph nodes.
• Develop automated ETL/ELT pipelines to map and integrate internal supply chain data with external, open-source ontologies and registries (such as GLEIF for corporate ownership, W3C SSN/SOSA for IoT sensors, and Copernicus for geo-hazard alerts, PROV-O for data provenance).
• Collaborate with AI/ML Engineers to establish low-latency GraphRAG retrieval layers—crafting optimized Cypher and SPARQL queries, implementing NL2Query tools for agents, hybrid vector-graph indexing pipelines, and Model Context Protocol (MCP) tool endpoints for autonomous LLM agents.
• Operationalize SHACL (Shapes Constraint Language) shapes into automated data quality tests within CI/CD pipelines to prevent hallucinated or non-compliant data mutations from entering the enterprise knowledge graph.
• Enhance multi-hop query performance, graph partitioning, and database indexing strategies to manage sub-second traversal across billions of nodes and edges.
• A degree in Computer Science, Mathematics, Engineering, or a related technical field.
• Over 4 years of practical experience in building and querying graph databases, specifically Labeled Property Graphs (Neo4j, Memgraph, TigerGraph) or RDF Triple Stores (GraphDB, Stardog, Virtuoso).
• Extensive experience with cloud technologies, preferably Azure and its ecosystem (e.g., Azure Foundry, Azure Bicep, AzureML, and Azure Cloud Storage).
• Advanced proficiency in Python (RDFLib, NetworkX, PyGraphistry) for developing scalable, production-quality data pipelines.
• Experience in creating entity extraction pipelines using modern NLP frameworks (LangChain, LlamaIndex, spaCy) or LLM-based structured extraction.
• Practical experience with contemporary data transformation tools (dbt) and integrating graph databases with vector stores (Qdrant, Pinecone, pgvector) for hybrid search architectures.
• Strong understanding of semantic web standards (RDF, RDFS, and OWL, SKOS, SHACL, RDF-star, SPARQL), graph schema design principles (T-Box vs. A-Box separation), and mapping languages for managing heterogeneous data structures (RML, R2RML).
• Experience with domain-specific supply chain, carbon accounting (GHG Protocol), or lifecycle assessment (LCA) data structures is a plus.
• Direct experience in building Model Context Protocol (MCP) servers to expose graph tools to LLM agents is a plus.
• Familiarity with enterprise OBDA approaches at scale is a plus.
• Support with all the necessary office and IT equipment.
• Flexible working hours.
• Wellness allowance for mental and physical wellbeing.
• Access to professional mental health support.
• Referral bonus policy.
• Opportunities for learning and development.
• Participation in sustainability events and community involvement.
• Peer recognition program.
• Employee-led resource groups.
• Optional (fully covered or co-financed) health care and life insurance.
• Multisport card.
• Multikafeteria.
• Lunch card.
• Hybrid work organization.
• Remote work from abroad policy.
• Internet and electricity bill allowance.
• Additional day for community service when volunteering.
LiteLLM AI Gateway
Snowflake
RTX
C-MORE
Get handpicked remote jobs straight to your inbox weekly.