RagLeap Docs
Docs for ragleap-core v0.7.5 · latest
Use

Knowledge graph

RagLeap Core builds a lightweight entity co-occurrence graph alongside its vector index. When you ingest a document, entities (product names, acronyms, proper nouns) are extracted and linked in Neo4j. When you ask a question, the same extraction runs on your query, and any documents linked to matching entities get a similarity boost in retrieval — on top of, not instead of, normal pgvector search.

Runs as a fourth Docker Compose service on ports 7475/7688 (remapped from Neo4j's defaults to avoid colliding with another Neo4j instance on the same host). If Neo4j is unreachable or NEO4J_PASSWORD is unset, the graph degrades gracefully — retrieval falls back to pure vector search, ingestion is unaffected.

Setup:

  1. Set NEO4J_URI, NEO4J_USER, and NEO4J_PASSWORD in .env (matching the NEO4J_AUTH value in docker-compose.yml)
  2. Optionally set DOMAIN_TERMS — a comma-separated list of domain-specific terms to boost during extraction (e.g. DOMAIN_TERMS=API,SDK,RAG)

Honest status: entity extraction, document graph writes, entity-based document lookup, and graph-boosted chat retrieval are all verified working end-to-end, including in CI (fresh build, real ingest, real query, real graph lookup). The graph boost is currently a simple additive score bump, not a full weighted re-ranker — a richer hybrid ranking system is a good next step for anyone who wants to dig in.

Known limitations:

  • Entity extraction is regex-based (CamelCase, acronyms, capitalized phrases, plus optional domain terms) — not a trained NER model, so it will miss some entities and occasionally include noise
  • search_related_entities() (multi-hop graph traversal) is implemented but not yet wired into the retrieval pipeline — good-first-issue candidate for anyone wanting a project

Generated from the ragleap-core v0.7.5 source. The repository is the source of truth and may be newer.