What it can learn from
| Source | What it reads |
|---|---|
| Your website | Every page the crawler reaches by following your links, or exactly the pages in your sitemap, or one page at a time |
| PDFs on your site | The PDFs your pages link to, read with the crawl, so the agent can link to them |
| Files | PDF, Word (DOCX), TXT and Markdown files you upload |
| Text | Anything you type or paste: opening hours, a policy, a price list |
| Q&A | A question and the exact answer you want given to it |
You do not need to prepare anything. The text is taken out, cut into passages and indexed for search, usually in seconds. A crawl obeys your robots.txt, reads one page at a time and does not run JavaScript, so a page whose text only appears after a script runs is read without it. Websites and Files have the details.
Training without fine-tuning
"Training" here does not change the model. When a visitor asks something, the agent searches your sources for the passages that answer it, by meaning and by the exact words (a product code, a name), and the model writes the answer from those passages only. That has three consequences that matter:
- Changes count at once. Edit a page and read it again, and the next answer uses the new text.
- It does not invent. When your content does not hold the answer, the agent says so, or asks what the visitor means, rather than guessing.
- You can check every answer. In Activity and the Playground, each one shows the pages it drew on. Your visitors see the answer, with a link when it comes from a PDF on your site.
On the Agency plan your site is read again every 7 days by itself, so prices, opening hours and new PDFs reach the agent without you thinking of it.
Knowing what it does not know
Every answer carries a confidence score: how sure the search was that your content held the answer, measured before a word is written. Confident answers also get a grounding check, a second opinion that reads the answer next to the passages and says whether they support it. The answers that scored low are listed in the Overview with what was missing, and from that list you add the right answer as a Q&A pair. That list is the fastest way to a better chatbot: it is a list of what your visitors asked and your content did not say. Fixing weak answers.
Why not ChatGPT on your own data?
A custom GPT lives inside ChatGPT: you cannot put it on your own website, and its users need a ChatGPT account. Building your own chatbot on a model's API is possible, and means writing the search, the widget, the language handling and the logging yourself. ChatterLab is that work done: your sources in, a widget out, with the sources and confidence of every answer in front of you, and the data in the EU.
Questions about training a chatbot
How long does it take to train a chatbot on my website?
A few minutes for most sites. The crawler reads one page at a time with a short pause between pages, so a site of a few hundred pages takes a little longer. Files are usually ready in seconds.
Can it read scanned PDFs?
No. A scan is a picture of text, and sources are not read by text recognition. Such a file is marked as failed. Run it through OCR first, or paste the text in as a text source.
Do I need to write code?
No. Adding sources, testing and styling happen in the dashboard. Putting the chat on your site is one script tag you copy and paste.
Is my data used to train AI models?
No. We do not train models on your content or your visitors' conversations, and OpenAI, which turns your text into search vectors, does not train on it either.
What if my content is in another language than my visitors?
The search works on meaning, so a French question finds a Dutch page, and the agent answers in French. For a market that matters, sources in its language still give the best answers.