Contents

Files from visitors

Let visitors send a document or a picture with their question, what the agent does with it, what it costs and how long it is kept.

With Visitors can send files switched on, the widget shows a paperclip beside the message box. A visitor attaches a document, or a picture on a model that reads images, and the agent reads it before it answers: "is this invoice right?", "what does my contract say about notice?", "what is broken here?". It is off until you switch it on, and it comes with Studio and up.

Switching it on

Under Settings, in Content, switch on Visitors can send files and save. Each agent has its own switch. On Free and Starter the switch is greyed out. If you move down from Studio, files stop being taken, and the switch keeps its setting for when you move back up. You can try files in the Playground first, with the switch still off.

What can be sent

KindFiles
PDFWith text you can select, or a scan, up to 20 pages
Word.docx
Excel.xlsx
PowerPoint.pptx
OpenDocument.odt, .ods, .odp
Plain text.txt, .md, .csv
PictureA photo or a screenshot, on a model that reads images

A Google Doc, Sheet or Slides file is sent by downloading it first as one of these. What a file is, is decided by its contents and not by its name: a renamed file is read as what it really is, or refused.

Not accepted, each with a sentence that tells the visitor why:

  • Pictures, on a model that reads no images. The visitor is asked to send the text or a document instead.
  • Old Office files (.doc, .xls, .ppt) and files with macros. The visitor is asked to save the file as .docx, .xlsx or .pptx.
  • Anything else, such as archives and web pages.

Scans

A PDF made of pictures of pages, a scan or a photographed letter, has no text for the reader to find. It is read by text recognition (OCR) instead, while it is uploaded: through Cortecs, like every answer, and only within the EU. Each page read that way costs 1 credit on top of the file. An agent reads at most 25 scans a day. Past that, or when text recognition finds no text either, the scan is refused with a sentence that says so. Text recognition can misread a word or a figure, and the agent is told that the text came from a scan.

Limits

LimitValue
Size of one file4 MB
Files with one message3
Files in one conversation10
Pages of a PDF20
Files from one visitor20 an hour, 50 a day
Files for one agent500 a day

A refused file counts toward the limits per visitor and per agent too: reading it is the work they protect.

What the agent reads

The agent reads a file together with the question it came with, and answers from both, and from your sources. All files of one message share about 4,000 tokens of text. A longer file is read from the start as far as that goes. The visitor then sees Only the first part is read under the file, and the agent is told that it saw only part.

A file is read with its own message only. In a follow-up question the agent knows that a file was sent and what it answered, not the file itself. A visitor who wants something else from it sends it again.

What a file says is material, never instructions. A document that tells the agent to behave differently is shown to it as the visitor's material, the agent is told not to follow it, and the audit log notes that the file read like instructions.

Pictures

A picture is taken only when the agent's model looks at images. In the list under Model, those models carry a picture mark. On any other model the paperclip offers documents only, and the switch says so. Models and credits.

A visitor attaches a picture with the paperclip, pastes a screenshot into the message box, or drops a file onto the chat. Before it is sent, the widget redraws it at 1,400 pixels on its longest side at most: a phone's photo then fits the size limit, and a screenshot of a form stays legible. Redrawing also leaves out what the camera stored in the file, such as where the photo was taken. The server takes that out a second time, for a picture that did not come through the widget.

A picture that says it is more than 30 megapixels is refused whatever the file weighs: what has to be unpacked is the pixels, by the browser of whoever opens the conversation and by the model.

The agent looks at the picture together with the question. Text it can read in a picture is the visitor's material, like a file's, and never an instruction. Unlike a document's text, it is not checked for attempts at instructions, because nothing reads it before the model does. A picture costs what a document costs.

What it costs

A file costs 2 credits on top of the answer on nearly every model, and more on the dearest ones: twice what reading a typical file costs. A page of a scan read by text recognition costs 1 credit more. A refused file costs nothing, and neither does one that was attached and never sent. Models and credits.

Where you see them

  • In Activity, under the visitor's message: the name, a small copy of a picture, the number of pages, scan when it was read by text recognition, and read in part when the agent read only the start. Click the name to download the file. Delete file erases it and what was read from it, and leaves the rest of the conversation. Activity.
  • In the export, as a line under the message.
  • In a client's monthly report by name only, with nothing to open.
  • Over the API, as attachments on the question.

How long files are kept

A file, and the text read from it, is kept for 30 days. After that the name stays in the conversation and the file is gone. A file that was attached and never sent is removed after 2 hours. Deleting a conversation deletes its files at once.

Files are stored in the EU, with Cloudflare R2. The text read from a file goes to the model with the question, through Cortecs, like the rest of the conversation. A picture goes the same way, inside the request itself and never as a link, so no provider fetches it from our storage. It is not sent to OpenAI: only the question is turned into search vectors. A follow-up question can repeat something the agent said about a file, and that question is. Data and privacy.