Every legal AI demo starts in the middle of the story: a lawyer asks a question, the platform answers with citations to the record. What the demo skips is the part that makes it possible - the processing that happens between "upload" and "ask". A scanned annexure cannot be cited until it has a text layer. A 200-file data room cannot be navigated until someone - or something - has worked out what each file is. That work is the document pipeline, and understanding it changes how well you use the tools built on top of it. This explainer walks Judicio's File Library pipeline stage by stage: OCR, entity extraction, summarisation, and organisation - and what each stage does and does not guarantee.
The pipeline nobody sees
When a file lands in the File Library - uploaded directly, imported from Google Drive, OneDrive, or SharePoint, or filed from an email - it goes through a multi-step pipeline before it is "ready": OCR if the document is scanned or image-based, entity extraction across the text, AI summary generation, and placement into the library's structure. The library accepts 25+ formats including ZIP bundles, so an entire matter's documents can arrive at once and queue through processing together.
The reason to care about the stages individually is that each one answers a different trust question. "Can I search this?" is an OCR question. "Why did the platform tag Acme Corp as a party?" is an extraction question. "Can I rely on this one-paragraph description?" is a summarisation question - and it has a different answer than the other two.
Two mechanics are worth knowing before the stages themselves. Duplicates are detected and skipped - re-uploading the same agreement in three matter folders does not process (or bill) it three times, which matters more than it sounds in due-diligence work where the same document arrives from every direction. And the cost model follows the pipeline: plans include a monthly allowance of upload pages processed free, and past it processing bills per page - uploads themselves are never blocked, and the upload screen shows the estimate before anything runs.
Stage 1: OCR - making scans into text
Optical character recognition converts scanned pages, faxes, photographed exhibits, and image-based PDFs into machine-readable text. In Indian and Gulf practice especially, a working matter file is full of these: stamped agreements scanned at a registry, certified copies, annexures photocopied through three generations. Until OCR runs, such a document is a picture - invisible to search, unusable for citation, unreadable by any AI feature.
The File Library runs OCR automatically on upload; there is no manual "make searchable" step. Practical implications follow. Full-text search covers the content of scans, not just filenames - so "arbitration" finds the clause even in the photocopied annexure. Quality of the source scan still matters: OCR on a clean scan is near-perfect, while a skewed, low-resolution photocopy can produce recognition errors - which is one reason downstream features cite to pages rather than silently paraphrasing, so the source is always one click away.
Stage 2: entity extraction - tagging what matters
Once text exists, extraction identifies and tags the structured facts inside it: parties, dates, monetary values, defined terms, and key provisions. This is the stage that turns a pile of text files into something a legal team can slice: every document mentioning a counterparty, every instrument dated before the transaction closed, every file with an indemnity provision.
| Entity type | Examples | What it enables |
|---|---|---|
| Parties | Companies, individuals, signatories | Party-wise filtering across a data room |
| Dates | Execution dates, deadlines, notice periods | Chronologies and deadline surfacing |
| Monetary values | Consideration, caps, penalties | Value-based triage of contracts |
| Defined terms | "Confidential Information", "Change of Control" | Consistency checks across documents |
| Key provisions | Indemnity, termination, governing law | Clause-level navigation without opening files |
Extraction is also what gives downstream features their hooks: a timeline is built from extracted dates and events; a bulk extraction workflow is extraction pointed at a specific question set. When you later ask "which of these agreements has a change-of-control clause", you are querying work that was done at upload time.
Stage 3: summaries - the navigable library
Every processed document receives a one-paragraph executive summary and a more detailed analysis. The honest way to describe their role: they are a navigation layer, not a substitute for reading. A summary answers "is this the supplemental deed or the original?", "which of these five leases covers the Pune premises?", "what is this 40-page annexure even about?" - the questions that otherwise cost an open-skim-close cycle per file.
At matter scale that navigation layer compounds. A due-diligence folder of 300 documents becomes a browsable catalogue where each entry announces itself; a litigation bundle's unnumbered annexures become findable. The professional discipline stays what it always was: for any document you rely on, the summary tells you where to look and the document remains the source - the same source-first principle behind AI document summarisation generally.
Stage 4: organisation - folders that build themselves
The final stage is structure. Auto-Folderise proposes a folder organisation for the files it has just read - by document type, party, or transaction phase - so a bulk upload becomes an organised workspace in seconds rather than an afternoon of drag-and-drop. Smart folders keep that structure alive as new files arrive. We covered the organisational workflow in depth in Taming Matter Files; the point here is its position in the pipeline: organisation is only possible because the previous stages produced text, entities, and summaries to organise by.
Everything is shareable under granular permissions - folders with team members or external parties - and full-text search runs across the entire library: OCR'd content, extracted entities, and summaries alike. Search that spans all three layers behaves differently from filename search in ways that change daily practice: a query can match a phrase inside a scanned annexure, a tagged party name, or a summary's description of a document whose own text never uses your search term. The first time "termination for convenience" surfaces a contract that phrases it as "cancellation without cause", the layered index has earned its keep.
What processing enables downstream
The pipeline is not an end in itself; it is the substrate every other feature stands on. A bulk document review can only check 80 contracts against a playbook because each contract is already text with known provisions. A review matrix can only answer 25 questions across a folder because the folder is already structured. Research grounded in your matter documents - and drafting that starts from a precedent in your library - both assume the library has actually read what it holds.
This is also the practical answer to "why not just keep files in a shared drive and paste into a chatbot?" A shared drive stores bytes; a processed library holds text, structure, and provenance. The difference shows up the first time you need to find one clause across a hundred scanned agreements - or prove where an AI answer came from.
The pipeline also feeds the surfaces outside the web app: emails filed from the Outlook and Gmail add-ins land in the same library and get the same processing, so an attachment filed from the inbox is as searchable and summarised as anything uploaded directly. One pipeline, every entry point.
The limits - and the professional checks that remain
Three limits are worth stating plainly. First, OCR fidelity depends on scan quality - a degraded photocopy can defeat any recognition system, so critical figures in poor scans deserve a human eye on the source page. Second, extraction is probabilistic: it is built to surface and tag, and the tags are hooks for navigation, not conclusions to file with a court. Third, summaries compress - by design they omit; anything you would stake a position on needs the underlying page, which is why every downstream output links back to sources rather than asking for trust.
Those limits are why the pipeline's job is best understood as preparation, not judgment. It does the reading a paralegal team would need days for, and leaves the professional decisions - what matters, what to rely on, what to file - where they belong.
Getting started with Judicio
The File Library is included in every Judicio plan: plans carry a monthly allowance of upload pages processed free, with per-page processing beyond it - and uploads are never blocked. The fastest way to understand the pipeline is to feed it a real bundle: start a free 7-day trial - 500 credits, no card required - upload a closed matter's folder, and watch what comes out the other side: searchable scans, tagged entities, a summarised, organised library ready for matrix review, timelines, and research.