What knowledge sources can an AI Twin ingest?

Every source type Reality Twin can ingest today — your website, uploaded documents (PDF, DOCX, Markdown, CSV and more) and pasted text — plus plan limits and what is coming next.

Updated June 18, 2026 6 min read

Your AI Twin only knows what you feed it. Reality Twin supports several source types, all indexed into one unified knowledge graph per twin. This guide covers every currently supported source, how they interact, and how to combine them for the most accurate answers.

Sources available today

Website crawl

The default source. Point at a root URL, we crawl and index — see Connect your website: how the AI Twin crawler works.

Documents (PDF, DOCX, MD, TXT, CSV)

Drag-and-drop upload for the material your website doesn't publish — spec sheets, price lists, policies, internal SOPs. Full details in Upload PDFs, DOCX and Markdown.

Structured Q&A

Canonical question-and-answer pairs written directly in the dashboard. They rank highest in retrieval, so use them to overrule ambiguous or contradictory sources.

Pasted text

Paste any text straight into the knowledge base — policies, FAQs, a price list, notes from a call. Useful for knowledge that lives in someone’s head rather than on a page.

Sitemap.xml import

For very large sites, supply a sitemap URL so no page is missed.

Plan limits

  • Free — website only, up to 100 pages.
  • Starter — website up to 1,000 pages + 10 uploaded documents + unlimited Q&A.
  • Pro — unlimited pages and documents.
  • Enterprise — all of the above plus Google Drive, Confluence, Zendesk (see below).

Coming soon

The following sources are on the near-term roadmap and available in private beta for Enterprise customers — talk to sales to opt in:

  • Google Drive (Docs, Sheets, Slides).
  • Confluence Cloud and Data Center.
  • Zendesk Help Center articles.
  • Salesforce Knowledge base.
  • Slack channels (opt-in per channel, private).

How the sources combine

All sources feed one knowledge graph. When two sources disagree, retrieval prefers, in order:

  1. Structured Q&A you authored directly (highest authority).
  2. Uploaded documents (assumed to be the most current internal reference).
  3. Website content (public, easiest to update).
Use structured Q&A as your override lever. When a customer keeps getting a subtly wrong answer, a single Q&A pair fixes it faster than editing pages.

Best practices

  • Prefer HTML over PDF where possible — HTML is easier to update and the crawler parses it more reliably.
  • Keep one canonical page per topic; delete duplicates.
  • Write one Q&A per hard question your team gets asked repeatedly.
  • Re-upload documents rather than adding new versions — the old file is replaced.

Frequently asked questions

Can I connect a database directly?

Not directly. Export the data to CSV or expose it as an HTML page and connect that. Direct database connectors are on the Enterprise roadmap.

How large a document library can I upload?

Pro workspaces have uploaded thousands of documents totalling several GB without issue. Per-file cap is 25 MB.

Are my sources used to train shared models?

No. Your sources are private to your workspace and never leave it for model training.

Did this article solve your problem?

If not, email us — a human on the founding team replies, usually within a business day.