RagLeap Core auto-detects language during document ingestion (per chunk)
and during chat (per query), using the langdetect library plus
script-based heuristics for CJK, Hangul, and Kana text. Since every
channel (WhatsApp, Telegram, Discord, Voice, and the API directly)
routes through the same core chat pipeline, detection applies
consistently everywhere without per-channel wiring.
Setup: works out of the box with no configuration. Optionally set
DEFAULT_LANGUAGE (fallback when detection fails or text is too short),
LANGUAGE_DETECTION_CONFIDENCE_THRESHOLD (default 0.7), and
LANGUAGE_DETECTION_SUPPORTED_LANGUAGES (comma-separated allowlist,
blank = unrestricted).
Honest status: verified working end-to-end — document-level detection tested at high confidence (0.9999) on a real mixed-language document, and query-level detection confirmed working via both the API and CLI.
Known limitations:
- Short queries in closely-related languages can be misdetected (in testing, a short French query was detected as Italian) — this is an inherent limitation of statistical detection on short text, not specific to this port. A good-first-issue candidate for anyone wanting to improve short-query accuracy
- Detection is one-way only: RagLeap Core detects the query's language and surfaces it, but does not yet steer the AI's response language to match — that's a reasonable next step for a contributor