Genetic.codes
All solutions
Searchable knowledge base from your documents

Thousands of documents nobody can search

We take the archive you already have, digital files and scanned paper alike, and turn it into a knowledge base your people can search and ask questions of through a portal.

Who it is forFor companies whose institutional memory is sitting in folders and filing cabinets.
The problem

The answer is in there somewhere.

Twenty years of contracts, specifications, correspondence and reports. Some of it born digital, much of it scanned at some point by somebody who has since left. The information exists, but reaching it means knowing which folder to open.

Scanned pages are the worst of it. To a computer a scan is a picture of a page, so searching for a supplier name across the archive returns nothing at all, even though the name is printed on forty of them.

Search returns nothing on scanned material
Only long-serving staff know where things are
The same question gets researched from scratch every time

The OCR step is what makes the rest possible, so it is worth being precise about what happens there.

Collect01

The existing archive is imported as it is: folders, shares, mailboxes, disks of scans. Structure is preserved.

Network shareObject storageMail archive
Read02

Digital documents give up their text directly. Scans and photographs are OCR’d page by page.

OCRSL + EN
Layer03

For scans, the recognised text is embedded back into the same PDF at the coordinates it came from.

Index04

Everything is indexed together, with metadata: type, date, parties, and where the file originally lived.

Metadata index
Ask05

People search, filter and ask questions in the portal, and get answers with the source page shown alongside.

Single sign-onPortal
Answer06

The answer comes back with the passage it was drawn from and the source page open beside it, so it can be checked rather than taken on trust.

The text layer is coordinate-mapped, which is why a hit lands on the right line of the right page instead of somewhere in the document.

What you get

OCR with a coordinate-mapped text layer

A scan becomes a searchable PDF that still looks exactly like the scan, and matches are highlighted on the page itself.

One search across everything

Digital and scanned material sit in the same index. You do not have to know which one you are looking for.

Questions, not just keywords

Ask what a clause said or which supplier delivered under a contract, and get an answer with the passage it came from.

Answers you can verify

Every result opens the source document at the relevant page. Nothing is asserted without the page behind it.

Access that follows your structure

Departments, roles and restricted material can be separated so people see the archive they are entitled to.

The archive stays yours

The result is your documents, improved in place. You can export the searchable PDFs and take them elsewhere.

What it fits with
Scanned PDFs, photographs and TIFFOffice documents and email archivesNetwork shares and object storageSingle sign-on for staff accessExport as searchable PDFAPI for your own front end

Slovenian and English are both handled in OCR, including mixed documents.