Files
Which documents an agent can read, what happens to them, and the one kind of PDF that silently teaches it nothing.
Upload the documents you already have: manuals, price lists, policies, product sheets.
PDFs that your website links to do not need uploading: a crawl of the site reads them, and the agent can link to them. See Websites. Upload what is not on the site.
What can be uploaded
On Sources, under Files, drop or choose PDF, DOCX, TXT or MD files. There is no fixed maximum per file. What limits you is the agent's storage allowance, and the size of the file counts towards it. Limits and freshness.
A file travels straight from your browser to our file storage, which is kept under EU jurisdiction. Data and privacy says where everything lives.
What happens next
- Uploading. The file is on its way. An upload that never finishes is cleaned up after 1 hour. One that fails on its way to storage is removed at once, and the page says the file could not be sent.
- Pending, then Processing. The text is taken out of the file, cut into passages of about 500 tokens that overlap a little so no sentence is lost at a seam, and indexed for search.
- Ready. The agent can use it. This usually takes seconds, a long document a little longer.
If a file ends on Failed, the Sources page says why: the file could not be read (a password on a PDF is the usual cause), no text was found in it, reading it took more memory or time than a source may use (split it into smaller files), or something went wrong on our side, in which case Retrain agent tries it again. A failed source teaches the agent nothing until it is fixed.
Scanned PDFs
Warning. A PDF that is a scan, a photograph of pages, contains no text, only pictures of text. Sources are not read by text recognition (OCR). Such a file ends on Failed, with the message that no text was found in it.
To check: open the PDF and try to select a sentence with your mouse. If you cannot select text, neither can we. Run the file through OCR first (most PDF programs and scanners can), or paste the content in as a text source.
Getting good answers out of documents
- Text beats layout. Tables that depend on merged cells, text in images and multi-column brochures lose their meaning when reduced to plain text. If a price list matters, a simple table or a text source answers better than a designed PDF.
- One subject per document finds better than one large everything-document, because each passage carries the document's title with it into the search. A file called "Returns policy" helps every passage inside it. A file called "scan_0042" helps none.
- Remove what is out of date. An agent cannot know that the 2023 price list has been replaced. If both are there, it may quote either.
Over the API
The API and an assistant connected over MCP cannot upload a file, so they add one from an address instead. We fetch it as the crawler fetches a page, within 20 seconds and up to 50 MB, and tell what it is from the file itself rather than from its name. A file that is only on your computer is uploaded here.