The answer is in there somewhere.
Twenty years of contracts, specifications, correspondence and reports. Some of it born digital, much of it scanned at some point by somebody who has since left. The information exists, but reaching it means knowing which folder to open.
Scanned pages are the worst of it. To a computer a scan is a picture of a page, so searching for a supplier name across the archive returns nothing at all, even though the name is printed on forty of them.
The OCR step is what makes the rest possible, so it is worth being precise about what happens there.
The existing archive is imported as it is: folders, shares, mailboxes, disks of scans. Structure is preserved.
Digital documents give up their text directly. Scans and photographs are OCR’d page by page.
For scans, the recognised text is embedded back into the same PDF at the coordinates it came from.
Everything is indexed together, with metadata: type, date, parties, and where the file originally lived.
People search, filter and ask questions in the portal, and get answers with the source page shown alongside.
The answer comes back with the passage it was drawn from and the source page open beside it, so it can be checked rather than taken on trust.
The text layer is coordinate-mapped, which is why a hit lands on the right line of the right page instead of somewhere in the document.
What you get
A scan becomes a searchable PDF that still looks exactly like the scan, and matches are highlighted on the page itself.
Digital and scanned material sit in the same index. You do not have to know which one you are looking for.
Ask what a clause said or which supplier delivered under a contract, and get an answer with the passage it came from.
Every result opens the source document at the relevant page. Nothing is asserted without the page behind it.
Departments, roles and restricted material can be separated so people see the archive they are entitled to.
The result is your documents, improved in place. You can export the searchable PDFs and take them elsewhere.
Slovenian and English are both handled in OCR, including mixed documents.
Point us at the archive
Tell us roughly how much there is and how much of it is scanned. That is enough for us to describe what the portal would look like for you.