The document isn't lost. It just isn't readable.
It is a scan, or a phone photo — so search never finds it.
“We have it” is not “we can use it”
The gap between those two is the job nobody does: open the document, read it, and turn it into something other people can find. That job is always deferred, and later never arrives.
For Persian documents the gap is wider still: Persian numerals, tables that collapse in conversion, a page scanned at an angle. General-purpose tools give up on exactly these.
Reading documents
From a photograph to structured text
Each page is read — Persian, Arabic, English; printed or scanned — and becomes structured Markdown. If a page cannot be read, it is reported by its own number with its reason. Nothing disappears silently.
Then you validate the pages that matter. Model output is a draft, not the truth.
- Persian numerals inside financial tables
- Page-by-page human validation
- Failed pages reported by number, never hidden
- ## Statement of financial position
- | Account | 1403 | 1402 |
- | Cash and equivalents | 84,120 | 61,540 |
- | Trade receivables | 402,880 | 366,210 |
Page 12 validated by Sara
The knowledge base
And then you can ask it
The knowledge base is built from documents, feed history and the structure of the records themselves — not from files alone. So “what was the payment condition in last year's contract?” has an answer that cites its page.
You do not have to trust the answer; you can open the page and look.
- Hybrid search: keywords and meaning together
- Every answer carries a source you can open
- Works without a PDF too — text is a document as well
What did we agree on Alborz's working capital adjustment?
The normalised working capital peg was set at 364.9 B T, based on the twelve-month average and excluding the related-party receivable of 41.2 B T.
Two fair questions
Does it work on bad scans?
That is the case it was built for. But we do not promise a hundred percent: if a page cannot be read, you get its number and the reason so you can decide — which is more honest than an invented confidence score.Does our data have to leave?
In our cloud, the document goes to the model endpoint to be read. On-premises, the model can live inside your own network too — and then nothing leaves it at all.
Bring the worst scan you have
The one everybody avoids opening. Watching what it does with that is the fastest way to judge this.
The free plan includes a monthly document-reading allowance.