RagLeap Docs
Docs for ragleap-core v0.7.5 · latest
Use

Retrieval & reliability

Beyond the core RAG pipeline, /chat (and the underlying core.chat.ask()) support several controls aimed at production use: retrieval quality, response latency, provider reliability, and cost.

Hybrid search (dense + sparse). By default, retrieval combines pgvector cosine similarity with Postgres full-text search (tsvector/ GIN index), fused via Reciprocal Rank Fusion — catching both semantic matches and exact keyword/identifier matches a pure embedding search can miss. Pass hybrid=false to use dense-only retrieval instead (cheaper — one query instead of two).

Streaming. POST /chat/stream streams the answer as it's generated (text/plain, chunked transfer) instead of waiting for the full response. Implemented natively per provider (Gemini, Anthropic, and OpenAI-compatible each have different streaming APIs — all three are real, not one stubbed).

Provider fallback. Set LLM_FALLBACK_PROVIDERS (comma-separated) to automatically retry with backup providers if the primary fails — a rate limit, outage, or bad key on your primary provider doesn't have to mean a failed request. Each fallback needs its own API key configured normally. Streaming can only fall back before any text has been sent to the caller — a mid-stream failure surfaces as an error rather than silently switching providers and confusing the output.

Generation controls. temperature, system_prompt, and max_tokens are all real per-call parameters (not just env-var defaults) — build your own agent behavior on top of RagLeap's retrieval without forking the library.

Token usage & context budget. Every blocking /chat call returns real token usage (prompt_tokens, completion_tokens, total_tokens) pulled directly from the provider's response — not an estimate. Retrieved context is also trimmed to MAX_CONTEXT_CHARS (default 12000, roughly 4 characters per token for English text) before being sent, dropping the lowest-ranked chunks first, so you're not paying for more context than necessary. Set MAX_CONTEXT_CHARS=0 to disable trimming.

Honest status: hybrid search's RRF fusion math verified correct against hand calculation. Streaming verified working end-to-end for the default provider. Provider fallback verified with a real broken-primary test — deliberately invalid API key, confirmed fallback to a working secondary provider with a correct answer. Token usage and context trimming verified with real numbers: a 3-chunk retrieval trimmed to 1 chunk under a tight budget reduced actual prompt_tokens by 38% on the same live API.

Known limitations:

  • Token usage reporting is not available for streaming responses — each provider's streaming API surfaces usage differently, and doing all three correctly is separate, not-yet-done work
  • MAX_CONTEXT_CHARS is a character-count approximation (~4 chars/token for English), not an exact per-provider tokenizer count
  • Hybrid search hasn't been benchmarked for actual ranking-quality improvement on a multi-document corpus with genuinely conflicting dense vs. sparse rankings — only correctness (fusion math, tokenization of unusual identifiers) has been verified so far

Generated from the ragleap-core v0.7.5 source. The repository is the source of truth and may be newer.