Upload PDFs, DOCX, Markdown and CSV to your AI Twin
Fill the gaps your website leaves. Add product docs, policies, spec sheets and internal knowledge — parsed, chunked and indexed automatically.
Documents are how you teach your AI Twin the things your public website deliberately leaves out — detailed pricing, specification sheets, internal SOPs, refund policies, warranty terms, product comparisons. Reality Twin parses, chunks and embeds every uploaded document within seconds so it's ready for retrieval.
How to upload
- Open Knowledge → Sources → Upload documents.
- Drag files onto the drop zone or click to select. Multi-file uploads are supported.
- Wait for each file's status to turn Indexed (typically 5–30 seconds per file).
- Ask your twin a question the document answers to confirm it retrieves correctly.
Supported formats
- PDF — text-based PDFs are parsed natively; scanned images without OCR are skipped.
- DOCX — Microsoft Word 2007+.
- Markdown (.md) — best format for structured internal docs.
- Plain text (.txt).
- CSV — parsed as a table with headers; each row becomes a retrievable fact.
Keep your document library clean
- Prefer well-structured documents with clear headings and short paragraphs.
- Remove draft and stale versions — the twin cannot tell them apart from current ones.
- Re-upload the file with the same name to update; the old version is replaced, not merged.
- Tag documents with categories in Knowledge → Sources for easier browsing.
Working with PDFs specifically
Text-based vs. scanned
PDFs generated from Word, InDesign, LaTeX, or web-to-PDF tools are text-based and parse perfectly. Scanned PDFs (an image of a document) do not — run them through OCR first (Preview on macOS, Adobe Acrobat, ocrmypdf on Linux).
Multi-column layouts
The parser handles two-column layouts, tables and footnotes. Very complex layouts (multi-column with sidebars and floating boxes) may occasionally interleave text — export from the source as single-column if you notice retrieval quality drop.
Best practices
- Rename files clearly (pricing-2026-q3.pdf beats final-final-v7.pdf).
- One topic per file makes citations clearer to visitors.
- For price lists that change frequently, publish them as an HTML page instead and connect via the crawler.
- For legally sensitive documents (terms, DPA, privacy), use structured Q&A so you control the exact wording of every answer.
Common mistakes
- Uploading a PDF and forgetting to remove the older version — the twin may cite both.
- Uploading a 200 MB brand book — reduce to just the sections the twin needs.
- Assuming the twin reads images inside PDFs. It reads text only; caption important diagrams in the surrounding text.
Frequently asked questions
How many documents can I upload?
Starter — 10. Pro — unlimited. Enterprise — unlimited with dedicated ingestion pipeline for very large libraries.
How long does indexing take?
Under 30 seconds for a typical 10-page PDF. Large documents (200+ pages) may take a minute or two.
Are my documents encrypted at rest?
Yes — AES-256 at rest, TLS 1.3 in transit. See Where is my data stored?
Can I upload confidential documents?
Yes. Set the document visibility to private so it is used only for your workspace's conversations and never exposed in Twin-to-Twin exchanges.
- What knowledge sources can an AI Twin ingest?Every source type Reality Twin can ingest today — your website, uploaded documents (PDF, DOCX, Markdown, CSV and more) and pasted text — plus plan limits and what is coming next.
- Force a manual re-crawl of your websitePush new pages, updated pricing or fresh content into your AI Twin immediately — from the dashboard, the API, or a CMS webhook.
- Fix wrong or outdated AI Twin answersWhen your AI Twin says something wrong, the fix is almost always in the source. Diagnose fast and remediate cleanly.
Did this article solve your problem?
If not, email us — a human on the founding team replies, usually within a business day.