Draft — wording to be confirmed by Martin before this goes live.
Where your data goes
Your material stays yours. To do its work, though, this service hands parts of it to a small number of other companies — a model that writes a description, a service that reads text off a scan, the place your files actually sit. This page is our best effort to name every one of them — we correct it whenever we trace one we had missed — and it says what is sent, where the work happens, how long it is kept, and what agreement covers it.
Nothing here is a promise made on your behalf. Where we can point at the exact setting that sends your material somewhere, we say so. Where we are going on someone else's word, or on a note in a file, we say that too — and where nobody has checked, we say that plainest of all.
How to read the labels
Every entry below carries one of three labels, telling you how sure we are about where that company does its work.
Checked in the code
We can point at the line that sends it there.
From our written notes
It is written down — in our own notes or the company's published terms — but there is no setting of ours that proves it.
Not checked yet
Our best understanding, and nobody has confirmed it. Treat it as a question, not an answer.
Of the 28 entries on this page, 6 are traced to the code, 7 rest on written notes, and 15 have not been checked. Every unchecked one says what is missing. Checked on 5 September 2026.
The companies that handle your collection material, and the companies that handle your account.
20 touch your collection material, and 10 touch your account. A company that genuinely does both — today, Neon and Vercel's hosting — is listed once in each section below, not just where most of what it holds sits.
Your collection material
Describing your material
When the system writes a description, pulls out names and dates, answers a question in Explore, or works out what a collection is about, it asks a Google model to do it. Every one of those requests is aimed at Google's European data centres.
Google Gemini — the model that writes descriptions
Checked in the code
What we send
The picture, page or recording itself, any text already read off it, and the description settings for your workspace.
Why
Describing an item, drafting a catalogue entry, and picking out the people, places, dates and subjects in it.
Where it happens
Google's European data centres.eu
How we know
The model gemini-3.5-flash is mapped to the eu region group in src/lib/ai/vertex-auth.ts, and calls go to aiplatform.eu.rep.googleapis.com.
How long they keep it
Google answers the request and does not keep the material afterwards, and does not use it to train its models.
Agreement
Google Cloud's data processing terms, with European regions chosen. Paid use of the models is not used for training.
Worth knowing. Where we send it is ours to prove. What Google does with it afterwards rests on Google's published terms — we cannot see inside their service.
Google Gemini — the model used for the heavier jobs
Checked in the code
What we send
Descriptions, transcripts and catalogue text from a whole accession, collection or dossier, and what you type into Explore.
Why
Analysing a collection or a fonds, deep reads in a dossier, and Explore conversations.
Where it happens
Google's European data centres. If the stronger model is switched on instead, that one runs in Belgium.eu / europe-west1
How we know
The heavier jobs use the model returned by getProModel() in src/lib/ai/vertex-auth.ts, and the same model-to-region map sends it to Europe either way.
How long they keep it
Google answers the request and does not keep the material afterwards, and does not use it to train its models.
Agreement
Google Cloud's data processing terms, with European regions chosen.
Google Gemini — the quick model
Checked in the code
What we send
Short pieces of text: a question typed into help, or a single passage being weighed up.
Why
Answering questions in the help panel, and working out which parts of a long text matter most.
Where it happens
Google's European data centres.eu
How we know
The model gemini-3.1-flash-lite is mapped to the eu region group in src/lib/ai/vertex-auth.ts.
How long they keep it
Google answers the request and does not keep the material afterwards.
Agreement
Google Cloud's data processing terms, with European regions chosen.
Google Gemini — search fingerprints
Checked in the code
What we send
Your descriptions and transcripts, in short pieces.
Why
Turning text into the numbers that let you search by meaning rather than by exact words, and that put suggested homes on an accession.
Where it happens
Google's European data centres.eu
How we know
The model gemini-embedding-2 is mapped to the eu region group in src/lib/ai/vertex-auth.ts.
How long they keep it
Google answers the request and does not keep the text. The numbers it returns are stored in our own database.
Agreement
Google Cloud's data processing terms, with European regions chosen.
Google Cloud Storage — the holding area for sound, film and long documents
From our written notes
What we send
The audio, video or long document file, for as long as the model is reading it.
Why
A model cannot be handed a large file directly. The file is put into a store in Europe and the model is pointed at it there.
Where it happens
A Google storage bucket in Europe.archivers-gemini-uploads-eu
How we know
The bucket name comes from getVertexGcsBucket() in src/lib/ai/vertex-auth.ts. Its European location and the two-day clear-out are recorded in docs/ARCHITECTURE.md §3.
How long they keep it
Files are removed automatically after two days.
Agreement
Google Cloud's data processing terms. The bucket is ours, in our own Google project.
Worth knowing. The name says Europe and our notes say Europe, but the location and the two-day rule are settings on Google's side. Worth confirming in the Google console before this page is signed.
Reading the text off a page
Turning a scan or a photograph of a document into text you can search and correct is done by Mistral, a French company.
Mistral — reading text off scans and photographs
From our written notes
What we send
The page image or the PDF.
Why
Turning a scan into text you can search, correct and export. This is switched off on the free Community plan.
Where it happens
Mistral's own service, EU-hosted. Mistral is a French company and processes this in the European Union.api.mistral.ai (EU-hosted; the address itself carries no region string)
How we know
The address and the model name are in src/lib/ai/mistral.ts. Mistral offers no separate region setting to select — the EU hosting is Mistral's standard processing location, confirmed as our settled position (3 Sep 2026, Martin Broadhurst).
How long they keep it
Mistral's terms say nothing is kept once the answer comes back, and that material sent to their service is not used to train their models.
Agreement
Mistral's platform terms. As a French company inside the EU, no extra transfer paperwork is needed.
Mistral — sorting pictures and spotting names
From our written notes
What we send
The picture, and the text that has been read off a document.
Why
Deciding whether something is a photograph, a document or an object, spotting names and places in text, and checking a piece of writing before it is offered to you.
Where it happens
Mistral's own service, EU-hosted, as above.api.mistral.ai (EU-hosted; the address itself carries no region string)
How we know
The models are named in src/lib/ai/mistral.ts. EU hosting confirmed as our settled position (3 Sep 2026, Martin Broadhurst) — see the OCR entry above.
How long they keep it
Mistral's terms say nothing is kept once the answer comes back.
Agreement
Mistral's platform terms.
Turning speech into text
Transcribing a recording is a separate job from describing it, and it uses a different model.
Mistral Voxtral — speech to text
From our written notes
What we send
The sound from your recording.
Why
Writing out what is said, and marking who is speaking. You ask for this item by item; it does not happen on its own, and it is switched off on the free Community plan.
Where it happens
Mistral's own service, EU-hosted, as above.api.mistral.ai (EU-hosted; the address itself carries no region string)
How we know
The model voxtral-mini-2602 is pinned in src/lib/ai/mistral.ts. EU hosting confirmed as our settled position (3 Sep 2026, Martin Broadhurst) — see the OCR entry above.
How long they keep it
Mistral's terms say nothing is kept once the answer comes back.
Agreement
Mistral's platform terms.
Checking names, places and subjects against outside authority files
When a description names a person, a place, an organisation or a subject, the workspace can check that name against public authority files — the same ones other archives and libraries use — so a place is a fixed point other catalogues agree on, not just a string of text. This happens automatically while an item is being described (switched on by default; off on the free Community plan) and whenever an archivist looks one up by hand. What goes out each time is the name itself, being matched — never a picture, a transcript, or anything else about the record. A match, once found, is kept in our own database, so the same name is not looked up a second time.
Library of Congress — subject headings
Not checked yet
What we send
The subject term being matched.
Why
Finding the Library of Congress's own heading for a subject.
Where it happens
The Library of Congress's own service. Not confirmed where it is physically handled.
How we know
src/lib/vocabulary/sources/lcsh.ts calls https://id.loc.gov/authorities/subjects/suggest2, registered in src/lib/vocabulary/resolve.ts:38.
How long they keep it
A public lookup service, to its own schedule; we have not set one.
Agreement
None — a public API with no account of ours and no agreement, rather than a commercial relationship.
Worth knowing. Nobody has read the Library of Congress's own data-handling terms for this service; its US federal ownership is public knowledge, not a setting of ours.
The Getty — art, place, agent and object vocabularies
Not checked yet
What we send
The term being matched — an object type, a place name, a person or organisation's name, or the subject of an image.
Why
Four vocabularies, one company: the Art & Architecture Thesaurus, the Thesaurus of Geographic Names, the Union List of Artist Names, and the Cultural Objects Name Authority.
Where it happens
The Getty Research Institute's own service. Not confirmed where it is physically handled.
How we know
src/lib/vocabulary/sources/aat.ts, tgn.ts, ulan.ts and cona.ts all call https://vocab.getty.edu/sparql.json, registered in src/lib/vocabulary/resolve.ts:39-40,46-47.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. Nobody has read the Getty's own data-handling terms for this service; its US base of operations (Los Angeles) is public knowledge, not a setting of ours.
OCLC — FAST subject headings and the international names file
Not checked yet
What we send
The subject term, or the person or organisation's name, being matched.
Why
Two lookups, one company: FAST subject headings, and VIAF, which brings together how different national libraries each record the same person or body.
Where it happens
OCLC's own service. Not confirmed where it is physically handled.
How we know
src/lib/vocabulary/sources/fast.ts calls https://fast.oclc.org/searchfast/fastsuggest; src/lib/vocabulary/sources/viaf.ts calls https://viaf.org/viaf/AutoSuggest for both the persons and the corporate-name lookups, registered in src/lib/vocabulary/resolve.ts:41,43-44.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. Nobody has read OCLC's own data-handling terms for this service; its US base of operations (Dublin, Ohio) is public knowledge, not a setting of ours.
Wikidata — the open, general-purpose authority file
Not checked yet
What we send
The term being matched — a name, a place or a subject.
Why
Finding an entry for something none of the specialist lists above cover.
Where it happens
The Wikimedia Foundation's own service. Not confirmed where it is physically handled.
How we know
src/lib/vocabulary/sources/wikidata.ts calls https://www.wikidata.org/w/api.php, registered in src/lib/vocabulary/resolve.ts:42.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. Nobody has read the Wikimedia Foundation's own data-handling terms for this service.
GeoNames — place names
Not checked yet
What we send
The place name being matched.
Why
Finding a fixed point for a place named in a description.
Where it happens
Not confirmed.
How we know
src/lib/vocabulary/sources/geonames.ts calls http://api.geonames.org/searchJSON, registered in src/lib/vocabulary/resolve.ts:45.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. GeoNames's own base of operations has not been confirmed.
Iconclass — subjects depicted in images
Not checked yet
What we send
The term being matched, describing what a picture depicts.
Why
Finding the Iconclass code for the subject of an image.
Where it happens
Not confirmed.
How we know
src/lib/vocabulary/sources/iconclass.ts calls https://iconclass.org/json/searchkeys, registered in src/lib/vocabulary/resolve.ts:48.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. Iconclass's own base of operations has not been confirmed.
US National Library of Medicine — medical subject headings
Not checked yet
What we send
The subject term being matched.
Why
Finding the standard medical subject heading for something in a health-related record.
Where it happens
The National Library of Medicine's own service, part of the US National Institutes of Health, a US federal agency. Not confirmed where it is physically handled.
How we know
src/lib/vocabulary/sources/mesh.ts calls https://eutils.ncbi.nlm.nih.gov/entrez/eutils/, registered in src/lib/vocabulary/resolve.ts:49.
How long they keep it
A public lookup service, to its own schedule.
Agreement
None — a public API with no account of ours and no agreement.
Worth knowing. Nobody has read the NLM's own data-handling terms for this service; its US federal ownership is public knowledge, not a setting of ours.
Drawing a map in the browser
Two screens draw a map — the search page's location map, and the authorities page's map of named places — both signed-in only, neither shown to a member of the public. Drawing one needs a mapping library and picture tiles that are not ours; both are fetched straight into the signed-in person's own browser rather than through our servers, the same way any website's images are.
unpkg — the mapping library's files
Not checked yet
What we send
The request itself, carrying the browser's own network address, as any request for a file on the web does.
Why
Fetching Leaflet's stylesheet and its marker icons, which are not bundled into our own code.
Where it happens
Not confirmed — a public content delivery network, no account of ours.
How we know
src/components/dashboard/LocationMap.tsx:33,40-42 and src/components/authorities/AuthorityMap.tsx:43 both fetch https://unpkg.com/leaflet@1.9.4/... directly in the browser.
How long they keep it
Not ours to set — a public file host, to its own schedule.
Agreement
None — a public file host, not a commercial relationship.
Worth knowing. unpkg's own base of operations and its logging practice have not been confirmed.
CARTO — the map tiles themselves
Not checked yet
What we send
Which small square of the map is being looked at, as coordinates in the request address — which, over the course of drawing the whole map, traces the shape of the archive's storage locations or its catalogued places.
Why
Drawing the grey background map underneath the pins.
Where it happens
Not confirmed.
How we know
src/components/dashboard/LocationMap.tsx and src/components/authorities/AuthorityMap.tsx:55 both request tiles from https://{s}.basemaps.cartocdn.com/light_all/{z}/{x}/{y}{r}.png directly in the browser.
How long they keep it
Not ours to set.
Agreement
None — a public tile service, not a commercial relationship.
Worth knowing. CARTO's own base of operations and its logging practice have not been confirmed.
Where your files and your catalogue sit
Two places: the files you upload, and the catalogue you build around them.
Vercel — where your uploaded files sit
Not checked yet
What we send
Every file you upload, along with the smaller copies made for the screen and the export files you generate.
Why
Holding your material so you and your colleagues can open it.
Where it happens
Vercel's file store. The place is chosen in Vercel's own settings rather than in our code, so we cannot read it from here.
How we know
src/lib/storage/blob.ts stores files without naming a region. The 28-day figure below is the retention setting for the Community plan in src/lib/pricing/catalog.ts.
How long they keep it
Kept until you delete them, or until your account is closed. On the free Community plan, files are removed after 28 days.
Agreement
Vercel's data processing addendum.
Worth knowing. Where the file store physically sits needs checking in the Vercel account. Our older sub-processor note says only "nearest region", which is not specific enough to sign — the Vercel dashboard itself is the only remaining way to settle this, and nothing in this repository or in the tools available to us reaches it.
Neon — the database holding your catalogue
From our written notes
What we send
Your account details, every catalogue field, the descriptions the models write, transcripts, and the record of who did what.
Why
Being your catalogue.
Where it happens
London.eu-west-2
How we know
The address of the database we develop against is in the eu-west-2 region. The live database is described as London (eu-west-2) in our privacy policy, in the published Data Processing Agreement's sub-processor table, and on the Trust page — three separate published documents agreeing. An older internal note, docs/compliance/SUB_PROCESSORS.md, said Frankfurt; that note was wrong and has been corrected to point here instead.
How long they keep it
Kept while your account is open. Closing an account removes everything within 30 days.
Agreement
Neon's data processing terms. Inside the UK and EU, so no extra transfer paperwork is needed.
Worth knowing. This rests on our own published documents agreeing with each other, not on a region string in this repository — the database we develop against is the only place eu-west-2 appears as a literal setting. Grouped here under collection material because that is most of what it holds, but your account row (name, email, password hash) lives in the same database, not somewhere separate.
Running the site
The site itself, and the overnight jobs that tidy up and send reminders, run on Vercel.
Vercel — running the site and the scheduled jobs
Checked in the code
What we send
Everything you do passes through here on its way to and from your browser — a page load, a file upload, a sign-in, alike.
Why
Serving the pages, handling uploads, and running the scheduled jobs.
Where it happens
London.lhr1
How we know
The regions setting in vercel.json names lhr1, which is London.
How long they keep it
Vercel keeps its own record of requests for a short period. We have not set that period ourselves.
Agreement
Vercel's data processing addendum.
Worth knowing. Grouped here under collection material because uploads are the clearest example, but every request passes through the same pipeline — including a sign-in, which is account data.
Your account
Where your files and your catalogue sit
Two places: the files you upload, and the catalogue you build around them.
Neon — the database holding your catalogue
From our written notes
What we send
Your account details, every catalogue field, the descriptions the models write, transcripts, and the record of who did what.
Why
Being your catalogue.
Where it happens
London.eu-west-2
How we know
The address of the database we develop against is in the eu-west-2 region. The live database is described as London (eu-west-2) in our privacy policy, in the published Data Processing Agreement's sub-processor table, and on the Trust page — three separate published documents agreeing. An older internal note, docs/compliance/SUB_PROCESSORS.md, said Frankfurt; that note was wrong and has been corrected to point here instead.
How long they keep it
Kept while your account is open. Closing an account removes everything within 30 days.
Agreement
Neon's data processing terms. Inside the UK and EU, so no extra transfer paperwork is needed.
Worth knowing. This rests on our own published documents agreeing with each other, not on a region string in this repository — the database we develop against is the only place eu-west-2 appears as a literal setting. Grouped here under collection material because that is most of what it holds, but your account row (name, email, password hash) lives in the same database, not somewhere separate.
Running the site
The site itself, and the overnight jobs that tidy up and send reminders, run on Vercel.
Vercel — running the site and the scheduled jobs
Checked in the code
What we send
Everything you do passes through here on its way to and from your browser — a page load, a file upload, a sign-in, alike.
Why
Serving the pages, handling uploads, and running the scheduled jobs.
Where it happens
London.lhr1
How we know
The regions setting in vercel.json names lhr1, which is London.
How long they keep it
Vercel keeps its own record of requests for a short period. We have not set that period ourselves.
Agreement
Vercel's data processing addendum.
Worth knowing. Grouped here under collection material because uploads are the clearest example, but every request passes through the same pipeline — including a sign-in, which is account data.
Keeping the service working
When something breaks we want to know, and we want to know which pages are slow. Two services do that.
Sentry — error reports
Checked in the code
What we send
When something goes wrong: the fault itself, the page it happened on, and a recording of the screen with the text and anything typed into a box hidden.
Why
So we find out a page is broken without waiting to be told.
Where it happens
Sentry's European service, in Germany.ingest.de.sentry.io
How we know
The address is written into src/sentry.server.config.ts and src/instrumentation-client.ts. The setting that would attach names, email addresses and network addresses to a report is deliberately switched off.
How long they keep it
Sentry keeps reports to its own schedule. We have not set a shorter one.
Agreement
Sentry's data processing addendum, on their EU-hosted service.
Vercel Analytics and Speed Insights — page counts and loading times
Not checked yet
What we send
Which page was viewed and how quickly it loaded.
Why
Seeing which parts of the service are used and which are slow.
Where it happens
Vercel's service. No region is chosen by us.
How we know
Both @vercel/analytics and @vercel/speed-insights are switched on in src/app/layout.tsx (imported at the top of the file, rendered at :52-53). Nothing in our code narrows what they collect or where they send it.
How long they keep it
To Vercel's own schedule.
Agreement
Vercel's data processing addendum.
Worth knowing. Two things here are somebody else's word rather than ours. Nothing in our code picks a region, so where this is handled is Vercel's choice; and that it holds no personal detail rests on Vercel's description of the product, not on anything we can see.
Google reCAPTCHA — keeping robots off the sign-up form
Not checked yet
What we send
A score from Google about whether a sign-up looks like a person, and the network address it came from.
Why
Stopping automated sign-ups. It runs on the sign-up form only.
Where it happens
Google's service. No region is chosen.
How we know
src/lib/captcha/recaptcha.ts checks the score against Google's general address, which carries no region.
How long they keep it
To Google's own schedule.
Agreement
Google's terms, with standard contractual clauses and the UK addendum covering the transfer.
Worth knowing. This service has no Europe-only option that we have set. Until someone confirms otherwise, assume it is handled outside Europe.
Messages we send you
Email for invitations, reminders and password resets. Text messages only if you choose them for signing in.
Resend — the email service
Not checked yet
What we send
Your name, your email address, and the message itself.
Why
Sending invitations, verification links, password resets and the reminders you have asked for.
Where it happens
Resend's service. We do not choose a region in our code.
How we know
src/lib/email/ builds every message and hands it to Resend without naming a region.
How long they keep it
Resend keeps a record of messages sent, to its own schedule. We have not set that period.
Agreement
Resend's data processing terms.
Worth knowing. Resend is an American company that offers European handling. Nothing in our code picks one. Check the Resend account settings before this page is signed.
Twilio — text-message security codes
From our written notes
What we send
Your mobile number and the one-time code — only if you have chosen text messages as your second step at sign-in.
Why
Sending the code that confirms it is you signing in.
Where it happens
The United States. This one is not handled in Europe.
How we know
src/lib/sms/twilio.ts reads the region and the edge from the settings TWILIO_REGION and TWILIO_EDGE. Neither is set in production as of 3 September 2026, so the default American routing applies. Twilio's European region does not yet carry text-message codes, which is why neither is set.
How long they keep it
Twilio keeps a record of the check, to its own schedule. No archive material is ever sent to Twilio.
Agreement
Twilio's data protection addendum, with standard contractual clauses and the UK addendum covering the transfer.
Worth knowing. The authenticator-app option makes no outside call at all and keeps everything in Europe. If American handling is a problem for your institution, use that instead — an administrator can require it for everyone.
Keeping our own mailing list up to date
A separate step to the messages above: when you sign up, verify your address, sign in, or your plan changes, a note goes to the mailing-list tool we use for the newsletter and product updates. Never anything from inside a workspace — this is account data, not collection material.
VBOUT — our mailing-list and marketing tool
From our written notes
What we send
Your name, your email address, your plan, whether your email is verified, and when you signed up and last signed in. A second, separate path on our marketing website sends an enquiry form's name, email, phone number, role and company to the same company when somebody asks our website assistant a question — this one is not tied to a workspace at all, and may not even be a customer.
Why
Keeping our own mailing list up to date, and sending the product-news and offer emails that are not part of the Service itself.
Where it happens
The United States.
How we know
src/lib/vbout/client.ts calls https://api.vbout.com/1 directly (:1). src/lib/vbout/sync.ts builds what is sent on signup, on email verification, on sign-in and on a plan change (syncNewUser, called from auth.service.ts). The public enquiry path is src/app/api/vbout-add/route.ts.
How long they keep it
To VBOUT's own schedule. We have not set a shorter one.
Agreement
VBOUT's own terms of service.
Worth knowing. This is account data — who you are, not what you have catalogued. No file, no description and no transcript is ever sent here. VBOUT being US-based rests on the company's known base of operations, not on a setting of ours; its own retention period, and whether a data processing agreement is in place, have not been confirmed.
Asking us for help
When you use the in-app support form, what you type becomes a support conversation with the company we use to run our helpdesk.
GetGist — our helpdesk
Not checked yet
What we send
Your name and email address, your organisation's name and your role in it, which plan you are on, the page you were on, the accession you were looking at if any, and the message itself.
Why
Running our support inbox, so a conversation with us is kept in one place rather than lost in email. Your identity comes from your own signed-in session, never from anything you type.
Where it happens
The United States. Not confirmed where the conversation is actually handled.
How we know
src/lib/integrations/getgist.ts calls https://api.getgist.com directly, with no region setting. src/lib/services/support.service.ts:16-64 builds what is sent.
How long they keep it
To GetGist's own schedule. We have not set a shorter one.
Agreement
GetGist's own terms of service.
Worth knowing. GetGist being US-based rests on the company's known base of operations, not a setting of ours; whether a data processing agreement is in place has not been confirmed.
Taking payment
Only for workspaces on a paid plan.
Stripe — subscriptions and credit purchases
Not checked yet
What we send
Your name, your email address and your billing details. Card numbers go straight to Stripe and never reach us.
Why
Taking the subscription payment and one-off credit purchases.
Where it happens
Stripe's service. Stripe is an American company with European handling for European customers.
How we know
Payments are handled by Stripe's own hosted checkout; nothing in our code sets a region.
How long they keep it
Stripe keeps payment records for as long as the law requires it to.
Agreement
Stripe's data processing agreement, with standard contractual clauses covering the transfer.
Worth knowing. Where Stripe handles a European customer's details has not been confirmed against Stripe's own account settings.
What each part of the service sends
The list above is arranged by company. This one is arranged by what you are doing. Each entry says what leaves this service when you use that part of it, who receives it, what is kept afterwards — by them and by us — and whether any of it is used to train a model. Each was traced through the code that builds the request, not from a description of how the system is meant to work.
Describe — what happens when you add something
What leaves
The thing itself, and what your workspace has told us about how to describe it. A photograph or an object is shrunk first — anything over about 100 KB is resized so its longest edge is 1,600 pixels and saved again as a JPEG — and travels inside the request. A document, a recording or a film is copied whole into our own storage in Europe, unshrunk, and the model is pointed at it there. With it goes your workspace's description settings (the institution you told us about, the writing style, the language to write in, and the reparative-description stance if you have switched it on), the collection's description and the descriptions of the collections above it, anything you typed about the item, the condition notes, the fields of your data model with their help text, any columns from a spreadsheet you uploaded with the batch, and — for a document — the text already read off it, in full. Nothing is held back for being closed: a record marked closed or restricted is described exactly like a public one, and its access setting is applied to the answer afterwards, not to what is sent.
Who sees it
Google's models in Europe for objects, documents, recordings and film. Mistral for photographs, and for the passes that pick out names and places, match a value to your word lists, and check the description against the picture.
What is kept
The description, on the record, and in that record's own history of changes, so you can see what the model said before anybody edited it. The copy of a document, recording or film that was staged for the model is deleted as soon as the model has answered, with a two-day clear-out behind it in case that fails. We keep no copy of what we asked or of the words that came back: the cost line records the model, the number of tokens and the price, and nothing else.
Training
No. Neither company uses this to train its models — for Google on its terms for paid use, for Mistral on its published terms. Both are somebody else's written word rather than a setting of ours, which is what the label on each company above is telling you.
Read the text off a page — turning a scan into words
What leaves
For a PDF, a web address rather than the file: we hand the company a link to the document as it sits in our file store, and it fetches the document itself. That link is not private. Our file store has no private mode — the address is unguessable rather than protected, and anyone holding it can open the document. For a scanned image the page itself goes, inside the request. A multi-page TIFF is turned into a PDF here first. Nothing of yours goes alongside it: no instructions, no catalogue text, not even the file's name.
Who sees it
Mistral's own service, hosted in the European Union.
What is kept
The text, on the record: the whole transcript, the pieces of it page by page, a short excerpt, and the reader's own confidence in each word. If your workspace has search by meaning switched on, that text is then sent onward to be turned into search fingerprints. For a PDF there is nothing at Mistral's end to delete, because nothing was uploaded — only pointed at. The one case where a file is uploaded, a converted TIFF, is deleted straight after the answer.
Training
No, on Mistral's published terms. That is their written word rather than a setting of ours, which is what the label on their entry above is telling you.
The whole recording, exactly as it is, inside the request. If it is a film, the film goes — pictures as well as sound. We do not pull the soundtrack out first. Nothing else goes with it: no title, no description, and not even your file's name, which is replaced by a made-up one. You ask for this recording by recording; it never happens on its own, and it is not part of the Community plan.
Who sees it
Mistral's own service, hosted in the European Union.
What is kept
The words, on the record, in timed pieces, with speaker labels where the model could tell the voices apart, and the recording's length. There is nothing at Mistral's end to delete, because the sound was sent inside the request rather than uploaded first.
Training
No, on Mistral's published terms — their written word rather than a setting of ours, as the label on their entry above says.
The written record: the collection's name and description, and for each item its title, its file name, what kind of thing it is, its description and its catalogue fields — together with the full text of anything that has been transcribed or read off the page. Your question goes with it. The pictures, recordings and documents themselves never leave: Explore reads the words, not the originals. A record you have closed is left out before the question is even assembled.
Who sees it
Google's models, in Google's European data centres.
What is kept
Google holds a working copy for one hour so a follow-up question is cheaper, then it goes. We keep the assembled text on the conversation itself, for as long as the conversation exists — no clock ages it out — so the next question does not have to rebuild it, and we keep your questions and the answers until you delete the conversation. Deleting a conversation removes both, ours and theirs.
Training
No. Nothing sent to a model here is used to train it.
The written record only. No picture, recording or document itself is ever sent. What goes is the dossier's title and your own description of it; for each collection its title, reference code, description and dates; for each accession its name, description, the repository, the donor's name and the provenance notes; and for each item its title, file name, kind, date, creator, description and subjects. The first pass stops there. A deep read adds the full text of anything transcribed or read off the page, trimmed to fit the reading budget your plan allows and shared out between the items — a passage too short to be worth trimming is left out altogether rather than cut to a stub. A follow-up question sends the whole of the previous analysis back with it, including any note a colleague left on a passage and the name of whoever reviewed it. Nothing is held back for being closed: unlike Explore, a dossier reads a closed or restricted record exactly like a public one — its donor name and provenance notes included, and on a deep read its full text too.
Who sees it
Google's models, in Google's European data centres.
What is kept
The analysis on the dossier, one step of undo behind it, your follow-up questions and their answers, and a note of what was actually read for each item — in full, in passages, trimmed, or not at all. A follow-up also keeps the assembled text on the dossier itself, for as long as the dossier exists; no clock ages it out, and it is cleared when the dossier's items change, when a new analysis finishes, and when the dossier is deleted. Google holds a working copy for one hour, then it goes.
Training
No. Nothing sent to a model here is used to train it, on Google's terms for paid use — which is somebody else's written word rather than a setting of ours.
Convert a catalogue — turning a paper catalogue into records
What leaves
Two companies, at two moments. First the source: your photographed pages, assembled here into one document, handed over as a web address the reading company fetches — the same not-private link described under reading the text off a page. Then, page by page, a photograph of the page goes to Google, shrunk if it is over about 350 KB, and with it the text already read off that page with the reader's own confidence marks still in it, the last 600 characters of the page before so an entry running across a page join is not cut in half, your template's columns with their labels and examples, whatever you wrote in the run's notes about the material, your reference prefix, and your workspace's description settings. A PDF that is already typed text sends no picture at all — only its words.
Who sees it
Mistral for the reading, Google's models in Europe for turning the page into entries.
What is kept
The assembled document and every straightened page image stay in your workspace as ordinary files; only the raw originals you photographed are removed, once assembly has succeeded. The text read off the source is written onto that source file. The run keeps its own progress, its first look at every page, and what happened to each one. Every record that comes out carries the page it came from, the model that read it, and the version of the template that shaped it.
Training
No — Google's terms for paid use, and Mistral's published terms. Both are written words rather than settings of ours, as the labels on their entries above say.
Usually nothing at all: an ordinary search is answered inside our own database and never leaves it. If your workspace has search by meaning switched on — and it is off until somebody turns it on — then what you typed is sent to Google on every search to be turned into a list of numbers. Nothing about who is asking goes with it: no name, no workspace, no address. Separately, when a record is indexed, its title, description, the names picked out of it and an excerpt of its text are sent to be turned into numbers, and its longer text in pieces of about 1,200 characters. If picture search is switched on, the picture itself goes; if document search is switched on, the whole PDF does.
Who sees it
Google's models, in Google's European data centres.
What is kept
The numbers, in our own database, beside the pieces of text they came from. The numbers for a search phrase are kept so that the same search need not be paid for twice: that store holds the numbers and not what you typed, has no workspace attached to it, and a phrase drops out seven days after the last time anyone searched it — from the last search, not from the first.
Training
No. Nothing sent to a model here is used to train it, on Google's terms for paid use — somebody else's written word rather than a setting of ours.
Signing in with an authenticator app. The codes are worked out on your own phone and checked here. Nothing leaves this service, and no other company is involved. If where your data is handled matters to your institution, this is the option to choose over text messages.
Our own testing. When we compare one model against another to decide which to use, our internal tool can send sample text to other providers, including ones based in the United States. That runs on our own test material in our development system, not on your workspace — but we have not written that boundary down anywhere binding, so treat it as a question to ask us rather than a promise.