From file to cited answer

No magic, no black box: this is the pipeline every document goes through, and why answers can carry citations you can trust.

  1. 1 - Upload

    Files go directly to private storage over an authorized upload. Nothing is public; nothing is shared between organizations.

  2. 2 - Extraction

    Text is extracted with its structure - pages, headings, sections. Encrypted or unreadable files are rejected with a precise reason, never silently dropped.

  3. 3 - Structure

    Repeated headers and footers are removed; headings and page boundaries are preserved so every passage keeps its location in the original.

  4. 4 - Chunking

    Documents are split into overlapping passages that respect paragraphs and headings, so a clause is never cut in half mid-thought.

  5. 5 - Embedding

    Each passage is converted into a semantic index entry, alongside a classic keyword index. A document is “ready” only when every passage is indexed.

  6. 6 - Retrieval

    A question runs against both indexes - exact wording and meaning - filtered by your permissions before anything is retrieved, then the best passages are selected.

  7. 7 - Generation

    The answer model sees only the retrieved passages and your question, with instructions to answer strictly from that evidence - or to say the evidence is insufficient.

  8. 8 - Citation mapping

    Every citation is checked against the passages that were actually supplied. An answer that cannot be verified against its sources is discarded, not shown.