Scanned Documents Now Import Into Second Brain
Your answers are only as complete as what made it into the database. Anything that arrived as a scan stayed out, which put the gap exactly where the signed documents were.
Import runs OCR now. Scanned files go in as searchable text with everything else.
How it works
The Knowledge files import card checks your machine and shows a label. OCR ready means scanned PDFs get read during import. Drop them in with the rest. No separate step, no converting first.
The recovered text is treated like any other document. Split into ordered sections, stored as rows, tagged, every word indexed. A 60-page scanned lease behaves like a typed one. Ask about the termination clause and Claude pulls the sections that mention it.
Reading them costs you nothing in tokens
Recognition runs through the OCR built into your operating system. Nothing is uploaded to be read. The scan is processed on your machine and the text goes into your local database.
The alternative is handing each page to an AI as an image. A full page runs well over a thousand tokens before a word of the answer is written, and you pay it again every session, because nothing was stored. Two hundred pages of contracts is a real bill for the privilege of turning a folder into text.
Here the reading happens once, at import. What sits in the database afterward is text, and a query pulls the few sections that matched instead of the pages.
What your machine needs
macOS 10.15 Catalina or later
Windows 10 or later. Windows may need a language pack. Settings, then Time & language, then Language & region.
If your system does not qualify you see OCR unavailable. The rest of the app works as before.
Two limits worth knowing before you point this at an archive. Recognition covers the languages your system supports, and a page in a language your OS cannot read comes back as noise. Printed type on a clean scan lands past 99 percent of characters. Handwriting, faded fax paper and photos shot at an angle can produce errors.
International text stops arriving broken
Detection for non UTF-8 documents is better in this build. Windows-1251, Shift-JIS, the ISO-8859 family and others get identified more reliably at import, so the text goes in readable and comes back searchable.
Also in this release
Higher import limits.
Tags apply consistently across a batch and stay searchable next to the content.
An interface pass, plus better help texts.
Try it on your mixed folder
Update at brain.hexact.io. Free download, 7-day trial, no card.
Find the scanned documents you gave up on, the old client contracts, and drop the folder on the Import page. Confirm it says OCR ready, run the import, then ask Claude about a clause.
Migrating an archive, or running several businesses off one database? Book a call and I will tell you what to import first.



