Document intelligence · One module of the operating system

Forward it. Drop it. Photograph it.
It arrives understood.

The Doc Inbox is where every document a job produces lands and is read. Oryn™ classifies what it is, extracts the facts that matter with the page each one came from, names it to a convention, files it to the right project and folder, and works out whether it supersedes a drawing already on file. Nothing that touches money or a legal identifier is written to your records by the machine. Values arrive amber, as candidates, and turn green when a person commits them.

The Founding Builders Programme · Onboarding in small cohorts

01 / What it does

What this feature does

Document management is the module that turns the paperwork of a build into a record the system understands. The discipline itself, what a construction document set contains and why revision control is a risk exercise rather than a filing exercise, is covered in the construction documents reference and the drawing revision control guide. This page is about the machinery underneath the Doc Inbox.

There are three doors in. Drop a file anywhere in the application. Forward an email to your organisation’s Oryn address, where each attachment is ingested on its own so one bad file never sinks the rest of the email, and where the body of the email itself becomes a document when there is enough of it to be worth keeping. Or paste text from the command palette, which is how the SMS that confirmed a date or the email that changed a scope stops living in one person’s phone. Every door leads to the same pipeline. A PDF with a text layer, a scan with none, a photograph of a document, a spreadsheet export and pasted text are all read, by different paths, into the same shape.

Because the module sits inside one operating system, a document read here is the same record that accounts payable matches an invoice against, that a plan set for takeoff is drawn from, and that answers a question later without anyone opening a folder. The intelligence layer as a whole is described on Oryn, and the category question of what to look for when you are comparing products is answered in the construction document management software buyer’s guide.

02 / Why it matters

Why builders need it

Filing is a job nobody has time for

Documents arrive faster than anyone files them, so they sit in an inbox until they are needed and cannot be found. A system that reads and files on arrival removes the queue rather than shortening it.

A superseded drawing that is still live is a rebuild

Two current versions of the same drawing is not an inconvenience, it is a trade building the wrong thing. The rule worth enforcing is one current version per drawing, resolved the moment a new one lands.

A silent overwrite destroys the record

Software that quietly replaces the old plan with the new one erases what the job was built from. A revision has to be a chain you can walk back along, not a file that changed while you were not looking.

An extracted number you cannot check is worse than none

If a value pulled off a document does not tell you which page it came from, checking it costs more than typing it did. Every fact carries its page, and one click puts the document in front of you at that page.

Confidence is not permission

A model can be confidently wrong about a legal identifier. So money and legal fields never commit on a confidence score in this system, whatever the number says. A person decides, or nothing is written.

The document you need is four years away

The certificate, the stamped approval and the signed variation matter most when a defect claim arrives long after handover. Documents that were read on arrival are still findable then, because they were understood, not just stored.

03 / The VIABUILD way

How VIABUILD handles it

Deterministic rules run first and are free. A model is only asked about what the rules could not resolve, and it is never asked for permission.

The order the pipeline runs in is the whole design. A free rules pass classifies the document from its first page. The text layer is pulled and split into page and line segments. Then a deterministic ladder tries five rungs in order for every field the document type expects, structured paths into the file’s own metadata, labelled templates, regular expressions, validating parsers, and rules learned from your corrections. First hit wins, and a hit that will not normalise cleanly is not treated as a fact at all. Only what is left after that reaches a language model, and it reaches it as one batched call for the unresolved fields. A fully machine-readable document can be processed end to end with no model call whatsoever, and re-processing a document never re-buys facts it already has.

Scans and photographs have no text layer, so they take the vision path instead. Where a scan is too heavy to send whole, its pages are rendered to downscaled images and sent as labelled pages, capped, with a loud note when the cap truncates the set rather than a quiet gap. Scanned pages are then transcribed into searchable segments, which is what makes a photographed document findable by a word inside it later.

Then Oryn does the part builders actually feel. It summarises what the document is, who the parties are, and which reference, drawing number, title, revision and date it carries. It matches the document to a project from those signals and, when it is confident enough about exactly one job, files it there and records the reasons why. It names the document to a convention where the job number is read from your database rather than extracted from the page, so the one part of the name that must be right never depends on a read. It files the document into the folder its type belongs in, unless you have corrected that filing before, in which case your correction wins.

Finally it resolves the document against what is already on file. Identical bytes are a true duplicate and chain under the original with nothing else run. A newer revision supersedes the current one and the old one is retained as superseded rather than deleted. An older drawing arriving late files itself into the chain in the right place instead of unseating a newer one. Two documents carrying the same revision label are flagged as a possible duplicate rather than archived, because two halves of a set can legitimately share a label. Anything the comparator cannot order stays a suggestion for a person. Auto-action only fires on strong drawing-number or reference matches, never on a title alone, and a stamped approval set is treated as an approval artefact that neither supersedes nor gets superseded.

  • Three intake doors, one pipeline, all reading into the same shape
  • Deterministic ladder first, a model only for what is left
  • Every fact carries the page it came from, one click away
  • Amber is a candidate, green is committed by a person
  • Money and legal fields never auto-commit, whatever the confidence
  • Auto-named from the job number in your database, not from the page
  • One current version per drawing, superseded versions retained
  • Reviews with named reviewers, due dates and an append-only trail

04 / The workflow

How it runs, step by step

  1. 01

    It arrives

    Dropped into the app, forwarded to your Oryn email address, or pasted as text. Attachments are handled one by one, and the original file is written to storage once and never modified afterwards.

  2. 02

    It is queued and read

    The document joins a work queue that is drained on a schedule, so nothing is processed inside the click that uploaded it. The Activity log shows where each document is up to, queued, reading, extracted, filed, needs you, or failed.

  3. 03

    The free rules go first

    Classification by rules, then the deterministic ladder across every field the type expects. Anything a rule can resolve costs nothing and never reaches a model, which is what makes reading every document affordable rather than selective.

  4. 04

    A model is asked only about the gaps

    One batched call for the fields still unresolved, over the text for a normal PDF or over the pages for a scan. The model is confined to the document in front of it, and the facts it returns carry their page like any other.

  5. 05

    Oryn works out what it is and where it goes

    A summary, the parties, the reference, drawing number, title, revision and date. From those it matches the job, names the document to the convention, and files it to the folder its type belongs in or the folder your past corrections say it belongs in.

  6. 06

    The revision question is answered on filing

    Identical bytes chain as a duplicate. A newer revision supersedes the current one and the old is kept. An older arrival files in as superseded. An equal label is flagged. Anything unclear stays a suggestion. Nothing is deleted and every action is reversible.

  7. 07

    You confirm what matters

    Facts sit amber in the review pane with their page beside them. Committing one writes it to the job record, stamps the provenance, and turns it green. A document can also be sent to named reviewers with a due date, and Oryn chases them for you.

  8. 08

    It stays findable and it keeps teaching

    Words inside the document, including a transcribed scan, are searchable across every job. And every time you correct a value Oryn read, the correction is banked privately so a rule can be written from the pattern rather than the same mistake being made again.

05 / FAQ

Common questions.

Three ways, all of which end in the same pipeline. You can drop a file anywhere in the application. You can forward an email to your organisation’s Oryn address, where each attachment is ingested separately so one unreadable file does not stop the others, and where the email body itself becomes a document when there is enough substance in it to be worth keeping. Or you can paste text straight from the command palette, which is how a message that changed a scope becomes part of the job record instead of staying in someone’s phone. Supported inputs include PDFs with a text layer, scanned PDFs with none, photographs and screenshots of documents, spreadsheet and CSV exports, and pasted text.

No, and the order of operations is the reason. A free rules classifier runs first, then a deterministic ladder tries five rungs for each expected field, structured metadata paths, labelled templates, regular expressions, validating parsers, and rules learned from your own corrections. A language model is only asked about the fields none of those resolved, and it is asked once, in a single batched call. A document that is fully machine readable can be processed with no model call at all, and re-processing a document never buys the same facts twice. The design point is that reading every document has to be cheap enough to do to every document, not just the important ones.

The revision question is answered at the moment the document is filed, against the drawings already on the job. If the bytes are identical to something already there, it is a true duplicate and chains under the original. If the new drawing orders as newer than the one on file, it supersedes it, and the superseded version is retained rather than deleted. If an older drawing arrives late, it files itself into the chain as superseded instead of unseating the current one. If two carry the same revision label, that is flagged as a possible duplicate rather than archived, because two halves of a set can legitimately share a label. Anything the comparison cannot order stays a suggestion for a person to settle. Auto-action only fires on a strong drawing-number or reference match, never on a similar title alone, and a council-stamped approval set is treated as an approval artefact that neither supersedes nor is superseded.

Amber is a candidate, a value Oryn extracted that is waiting on a person. Green is committed, a value a human accepted. Committing a fact writes it into the job record, records the provenance of the change including where it came from and who accepted it, and flips the fact to committed. The rule underneath it is that fields carrying money or a legal identifier stay candidates structurally, regardless of how confident the extraction was. Confidence describes how cleanly a value was read, not whether it is correct, so it is never treated as authorisation. A high confidence score on a wrongly located approval number is exactly the failure this design exists to catch.

Naming follows a convention that starts with the job number, then the document label, then the revision where a revision is meaningful, then the document date. The job number is read from your own database rather than extracted from the page, so the part of the name that has to be right never depends on a read. Renaming only happens while the document still carries the filename it was uploaded with, so a name you chose is never overwritten. Filing goes to the folder the document type belongs in, within a standard project folder structure covering estimating, plans, specifications, engineering, supplier quotes, takeoffs, contracts, approvals, selections, construction, variations, claims, invoices, compliance and handover. If you have moved that kind of document before, your correction becomes the preference and beats the default.

Yes. A review is a specific document at a specific version, assigned to a named reviewer with a due date, and the project manager is pre-selected as the default reviewer. Each reviewer gets their own review and their own notification. The states are deliberately few, requested, then approved, approved as noted, rejected or cancelled, and nothing moves twice. A resubmission is a new review on the new version rather than a reopened old one, which keeps the history honest. Every transition writes an append-only event. If a review sits unanswered past your chase setting, Oryn chases it automatically, at most once a day, and never once a decision has been made.

Neither, and that is deliberate. Search runs over the words actually in your documents, including scanned pages that were transcribed, plus the document names and Oryn’s own summaries and tags. Results come back with the matching snippet and the page it sits on. What you get is recall grounded in your own material rather than a similarity score you cannot inspect, and when nothing relevant is on file the honest answer is that nothing is on file.

Yes. Automatic filing and automatic superseding are governed by the same organisation-level setting, and turning it off leaves Oryn suggesting rather than acting while the reading and extraction continue as normal. Everything the automation does is reversible in any case. Filing a document to the wrong project is fixed by reassigning it, and superseding a drawing flips a status and adds a forward link rather than removing anything, so the original document and every version of it remain.

It banks them, and a person promotes them. When you override a value Oryn extracted, the correction is recorded privately to your organisation. Nothing reads that bank automatically. An administrator can see which fields are being corrected most, write a rule against the real corrected examples, test it against them, and promote it only once it demonstrably extracts the right thing. A rule that matches none of the banked examples is refused. The pipeline only ever improves through a promoted rule, which means learning cannot quietly change what your documents say. Filing preferences work the same way, the folder you moved a document kind into becomes the preference for that kind next time.

06 / Keep reading

Related features & guides

Send it the messiest job folder you have.

Forward a week of job emails, drop in a plan set, photograph a stamped approval. Watch each one arrive named, filed against the right job, with the facts it carries citing the page behind them and waiting for you to confirm.