Upload PDFs, DOCX, Markdown and CSV to your AI Twin

Fill the gaps your website leaves. Add product docs, policies, spec sheets and internal knowledge — parsed, chunked and indexed automatically.

Updated June 19, 2026 6 min read

Documents are how you teach your AI Twin the things your public website deliberately leaves out — detailed pricing, specification sheets, internal SOPs, refund policies, warranty terms, product comparisons. Reality Twin parses, chunks and embeds every uploaded document within seconds so it's ready for retrieval.

How to upload

  1. Open Knowledge → Sources → Upload documents.
  2. Drag files onto the drop zone or click to select. Multi-file uploads are supported.
  3. Wait for each file's status to turn Indexed (typically 5–30 seconds per file).
  4. Ask your twin a question the document answers to confirm it retrieves correctly.

Supported formats

  • PDF — text-based PDFs are parsed natively; scanned images without OCR are skipped.
  • DOCX — Microsoft Word 2007+.
  • Markdown (.md) — best format for structured internal docs.
  • Plain text (.txt).
  • CSV — parsed as a table with headers; each row becomes a retrievable fact.
Maximum file size: 25 MB. For larger files, split into logical chunks (e.g. one PDF per product line) — the twin actually performs better with well-scoped documents.

Keep your document library clean

  • Prefer well-structured documents with clear headings and short paragraphs.
  • Remove draft and stale versions — the twin cannot tell them apart from current ones.
  • Re-upload the file with the same name to update; the old version is replaced, not merged.
  • Tag documents with categories in Knowledge → Sources for easier browsing.

Working with PDFs specifically

Text-based vs. scanned

PDFs generated from Word, InDesign, LaTeX, or web-to-PDF tools are text-based and parse perfectly. Scanned PDFs (an image of a document) do not — run them through OCR first (Preview on macOS, Adobe Acrobat, ocrmypdf on Linux).

Multi-column layouts

The parser handles two-column layouts, tables and footnotes. Very complex layouts (multi-column with sidebars and floating boxes) may occasionally interleave text — export from the source as single-column if you notice retrieval quality drop.

Best practices

  • Rename files clearly (pricing-2026-q3.pdf beats final-final-v7.pdf).
  • One topic per file makes citations clearer to visitors.
  • For price lists that change frequently, publish them as an HTML page instead and connect via the crawler.
  • For legally sensitive documents (terms, DPA, privacy), use structured Q&A so you control the exact wording of every answer.

Common mistakes

  • Uploading a PDF and forgetting to remove the older version — the twin may cite both.
  • Uploading a 200 MB brand book — reduce to just the sections the twin needs.
  • Assuming the twin reads images inside PDFs. It reads text only; caption important diagrams in the surrounding text.

Frequently asked questions

How many documents can I upload?

Starter — 10. Pro — unlimited. Enterprise — unlimited with dedicated ingestion pipeline for very large libraries.

How long does indexing take?

Under 30 seconds for a typical 10-page PDF. Large documents (200+ pages) may take a minute or two.

Are my documents encrypted at rest?

Yes — AES-256 at rest, TLS 1.3 in transit. See Where is my data stored?

Can I upload confidential documents?

Yes. Set the document visibility to private so it is used only for your workspace's conversations and never exposed in Twin-to-Twin exchanges.

Did this article solve your problem?

If not, email us — a human on the founding team replies, usually within a business day.