Beyond the core RAG pipeline, /chat (and the underlying core.chat.ask())
support several controls aimed at production use: retrieval quality,
response latency, provider reliability, and cost.
Hybrid search (dense + sparse). By default, retrieval combines
pgvector cosine similarity with Postgres full-text search (tsvector/
GIN index), fused via Reciprocal Rank Fusion — catching both semantic
matches and exact keyword/identifier matches a pure embedding search can
miss. Pass hybrid=false to use dense-only retrieval instead (cheaper —
one query instead of two).
Streaming. POST /chat/stream streams the answer as it's generated
(text/plain, chunked transfer) instead of waiting for the full response.
Implemented natively per provider (Gemini, Anthropic, and OpenAI-compatible
each have different streaming APIs — all three are real, not one stubbed).
Provider fallback. Set LLM_FALLBACK_PROVIDERS (comma-separated) to
automatically retry with backup providers if the primary fails — a rate
limit, outage, or bad key on your primary provider doesn't have to mean a
failed request. Each fallback needs its own API key configured normally.
Streaming can only fall back before any text has been sent to the
caller — a mid-stream failure surfaces as an error rather than silently
switching providers and confusing the output.
Generation controls. temperature, system_prompt, and max_tokens
are all real per-call parameters (not just env-var defaults) — build your
own agent behavior on top of RagLeap's retrieval without forking the
library.
Token usage & context budget. Every blocking /chat call returns real
token usage (prompt_tokens, completion_tokens, total_tokens) pulled
directly from the provider's response — not an estimate. Retrieved
context is also trimmed to MAX_CONTEXT_CHARS (default 12000, roughly
4 characters per token for English text) before being sent, dropping the
lowest-ranked chunks first, so you're not paying for more context than
necessary. Set MAX_CONTEXT_CHARS=0 to disable trimming.
Honest status: hybrid search's RRF fusion math verified correct
against hand calculation. Streaming verified working end-to-end for the
default provider. Provider fallback verified with a real broken-primary
test — deliberately invalid API key, confirmed fallback to a working
secondary provider with a correct answer. Token usage and context
trimming verified with real numbers: a 3-chunk retrieval trimmed to 1
chunk under a tight budget reduced actual prompt_tokens by 38% on the
same live API.
Known limitations:
- Token usage reporting is not available for streaming responses — each provider's streaming API surfaces usage differently, and doing all three correctly is separate, not-yet-done work
MAX_CONTEXT_CHARSis a character-count approximation (~4 chars/token for English), not an exact per-provider tokenizer count- Hybrid search hasn't been benchmarked for actual ranking-quality improvement on a multi-document corpus with genuinely conflicting dense vs. sparse rankings — only correctness (fusion math, tokenization of unusual identifiers) has been verified so far