RagLeap Core builds a lightweight entity co-occurrence graph alongside its vector index. When you ingest a document, entities (product names, acronyms, proper nouns) are extracted and linked in Neo4j. When you ask a question, the same extraction runs on your query, and any documents linked to matching entities get a similarity boost in retrieval — on top of, not instead of, normal pgvector search.
Runs as a fourth Docker Compose service on ports 7475/7688 (remapped
from Neo4j's defaults to avoid colliding with another Neo4j instance on the
same host). If Neo4j is unreachable or NEO4J_PASSWORD is unset, the graph
degrades gracefully — retrieval falls back to pure vector search, ingestion
is unaffected.
Setup:
- Set
NEO4J_URI,NEO4J_USER, andNEO4J_PASSWORDin.env(matching theNEO4J_AUTHvalue indocker-compose.yml) - Optionally set
DOMAIN_TERMS— a comma-separated list of domain-specific terms to boost during extraction (e.g.DOMAIN_TERMS=API,SDK,RAG)
Honest status: entity extraction, document graph writes, entity-based document lookup, and graph-boosted chat retrieval are all verified working end-to-end, including in CI (fresh build, real ingest, real query, real graph lookup). The graph boost is currently a simple additive score bump, not a full weighted re-ranker — a richer hybrid ranking system is a good next step for anyone who wants to dig in.
Known limitations:
- Entity extraction is regex-based (CamelCase, acronyms, capitalized phrases, plus optional domain terms) — not a trained NER model, so it will miss some entities and occasionally include noise
search_related_entities()(multi-hop graph traversal) is implemented but not yet wired into the retrieval pipeline — good-first-issue candidate for anyone wanting a project