What knowledge sources can an AI Twin ingest?
Every source type Reality Twin can ingest today — your website, uploaded documents (PDF, DOCX, Markdown, CSV and more) and pasted text — plus plan limits and what is coming next.
Your AI Twin only knows what you feed it. Reality Twin supports several source types, all indexed into one unified knowledge graph per twin. This guide covers every currently supported source, how they interact, and how to combine them for the most accurate answers.
Sources available today
Website crawl
The default source. Point at a root URL, we crawl and index — see Connect your website: how the AI Twin crawler works.
Documents (PDF, DOCX, MD, TXT, CSV)
Drag-and-drop upload for the material your website doesn't publish — spec sheets, price lists, policies, internal SOPs. Full details in Upload PDFs, DOCX and Markdown.
Structured Q&A
Canonical question-and-answer pairs written directly in the dashboard. They rank highest in retrieval, so use them to overrule ambiguous or contradictory sources.
Pasted text
Paste any text straight into the knowledge base — policies, FAQs, a price list, notes from a call. Useful for knowledge that lives in someone’s head rather than on a page.
Sitemap.xml import
For very large sites, supply a sitemap URL so no page is missed.
Plan limits
- Free — website only, up to 100 pages.
- Starter — website up to 1,000 pages + 10 uploaded documents + unlimited Q&A.
- Pro — unlimited pages and documents.
- Enterprise — all of the above plus Google Drive, Confluence, Zendesk (see below).
Coming soon
The following sources are on the near-term roadmap and available in private beta for Enterprise customers — talk to sales to opt in:
- Google Drive (Docs, Sheets, Slides).
- Confluence Cloud and Data Center.
- Zendesk Help Center articles.
- Salesforce Knowledge base.
- Slack channels (opt-in per channel, private).
How the sources combine
All sources feed one knowledge graph. When two sources disagree, retrieval prefers, in order:
- Structured Q&A you authored directly (highest authority).
- Uploaded documents (assumed to be the most current internal reference).
- Website content (public, easiest to update).
Best practices
- Prefer HTML over PDF where possible — HTML is easier to update and the crawler parses it more reliably.
- Keep one canonical page per topic; delete duplicates.
- Write one Q&A per hard question your team gets asked repeatedly.
- Re-upload documents rather than adding new versions — the old file is replaced.
Frequently asked questions
Can I connect a database directly?
Not directly. Export the data to CSV or expose it as an HTML page and connect that. Direct database connectors are on the Enterprise roadmap.
How large a document library can I upload?
Pro workspaces have uploaded thousands of documents totalling several GB without issue. Per-file cap is 25 MB.
Are my sources used to train shared models?
No. Your sources are private to your workspace and never leave it for model training.
- Upload PDFs, DOCX, Markdown and CSV to your AI TwinFill the gaps your website leaves. Add product docs, policies, spec sheets and internal knowledge — parsed, chunked and indexed automatically.
- Force a manual re-crawl of your websitePush new pages, updated pricing or fresh content into your AI Twin immediately — from the dashboard, the API, or a CMS webhook.
- Fix wrong or outdated AI Twin answersWhen your AI Twin says something wrong, the fix is almost always in the source. Diagnose fast and remediate cleanly.
Did this article solve your problem?
If not, email us — a human on the founding team replies, usually within a business day.