
GraphRAG Engineer
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in Spain, +2 more countries.
β’ Provision, optimize, and sustain production-quality Neo4j graph databases and pgvector vector storage clusters.
β’ Design high-throughput indexing structures, cosine similarity vector indexes, and query enhancements for sub-second response times.
β’ Create automated data ingestion pipelines that parse Git repositories, abstract syntax trees (ASTs), Jira issue links, Apache Avro schemas, and CI/CD metadata into an enterprise knowledge graph.
β’ Link distributed pipeline engines to hybrid retrievers that integrate SQL, Cypher graph traversals, and dense vector embeddings.
β’ Set up circuit breakers, confidence scoring thresholds, and step-limit constraints for autonomous agents.
β’ Integrate microservices and knowledge repositories with the Enterprise AI Gateway.
β’ Maintain version-controlled system prompt structures in localized .ai/ spoke directories while adhering to DLP PII scrubbing guidelines and token rate limits.
β’ Implement automated failover, backup restoration, and multi-cloud storage-tier cost management across AWS and GCP.
β’ Take ownership of the semantic, vector, and graph storage layer that drives the context engine for enterprise AI utilities and the Internal Developer Portal.
β’ Lead the deployment of the SDLC Context Graph and GraphRAG Engine for CAB compliance, code/schema lineage tracking, and enterprise LLM proxy integrations.
β’ Extensive operational and development experience with Neo4j (Cypher, APOC, causal clustering) or enterprise Knowledge Graphs.
β’ Demonstrated expertise with pgvector (PostgreSQL), embeddings management, hybrid search methodologies, and framework integrations (LangChain, LlamaIndex, or custom RAG pipelines).
β’ Practical experience managing relational (PostgreSQL) and graph databases within AWS and GCP cloud environments.
β’ Proficient in consuming Apache Avro payloads, streaming Kafka events (AWS MSK), and parsing both structured and unstructured code and JSON artifacts.
β’ Solid understanding of Prompts-as-Code patterns, few-shot prompt optimization, and tool specification for agents.
β’ Experience with Infrastructure-as-Code (Terraform) principles, Kubernetes (EKS/GKE), Docker, and pull-based GitOps workflows.
β’ Familiarity with HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities.
β’ Knowledge of OpenTelemetry (OTel) instrumentation for monitoring vector search query latencies and LLM inference performance in Datadog or Grafana.
β’ Generous PTO Policy.
β’ Unplugged Days supporting work-life balance.
β’ Flexible Work From Home Policy.
β’ Mental & Physical Wellness programs.
β’ Phone and Internet Reimbursement program.
β’ Access to Continued Career Development.
β’ Comprehensive Benefits and Competitive Packages.
β’ Paid Volunteer Time.
β’ Employee Resource Groups.
Green Energy Venture AG
Abacus Group
EXP
Doosan
Get handpicked remote jobs straight to your inbox weekly.