Skip to main content
All articles

Law Firm Operations

What actually breaks when a firm scans its own backfile

Most in-house digitisation projects do not fail on the scanning. They fail on separation, naming and the moment nobody can find the 2016 lease.

Written by
Helena Croft
Head of Capture Services, Datacoll8
Published
Reading time
9 minutes

A partner signs off the scanner, someone clears a store room, and for three weeks the project looks like it is working. Boxes empty, the counter on the device climbs, and the archive shrinks. Then a fee earner asks for the second amendment to a 2016 lease, and it takes forty minutes to find, because it was scanned into a two hundred page file called SCAN_0043.pdf along with four unrelated agreements.

That is the shape of almost every in-house digitisation project we are called into after the fact. The hardware was rarely the problem. The problem is that scanning was treated as a task rather than a process, and the decisions that determine whether the output is usable were never made by anybody.

Failure one: no separation strategy

A feeder does not know where one document ends and the next begins. Unless you tell it, a box of paper becomes one enormous file. The two workable answers are barcoded separator sheets inserted during preparation, or automatic separation driven by patch codes and cover pages. Both cost time at the front of the process and save a great deal at the back of it.

What does not work is asking operators to split files afterwards in a PDF editor. It is slow, it is unlogged, and it is the point at which pages start to end up in the wrong document. If you are choosing between preparing paper properly and correcting output later, prepare the paper.

Failure two: naming decided at the device

When each operator names files by hand, you get four conventions in a month and none of them survive a staff change. The name should be assembled by the system from indexed values: matter reference, document type, counterparty, date. If the value is not indexed, it cannot be in the name, and if it is in the name but not indexed, your search is relying on somebody having typed it correctly at four in the afternoon.

Failure three: image quality nobody set

Default driver settings are tuned for readable office documents, not for extraction. Greyscale scans of a faxed schedule at 200 dpi look fine to a human and read badly to a model. Deskew, despeckle, colour normalisation and a resolution appropriate to the paper are decisions, not preferences, and they should be locked into a profile so no operator has to make them twice.

Failure four: no verification step at all

If indexing is done by hand, someone is typing a matter reference hundreds of times a day, and a percentage of those will be wrong. If it is done by extraction with no threshold and no queue, wrong values are written into the record with the same authority as correct ones. Either way, the error is invisible until a fee earner needs the document.

The fix is a confidence threshold and a review queue. Values the system is sure about go straight through. Values it is unsure about are shown to a reviewer next to the page region they came from. That single arrangement converts a silent error rate into a visible workload you can measure and shrink.

The question is not whether errors happen. It is whether you find out about them during scanning or during a matter.
Helena Croft, Head of Capture Services

Failure five: the backfile eats the daily post

Bulk digitisation and daily inbound compete for the same device and the same people. Firms almost always start with the backfile because it is visible, and inbound post quietly builds a second, newer paper archive behind it. Sequence it the other way. Get today under control first, then work backwards through the store room, because a backfile is not getting any larger and today is.

What a working setup looks like

  • Separation decided before scanning starts, by barcode or cover sheet, never by hand afterwards
  • Filenames assembled from indexed values, not typed at the device
  • Imaging profiles locked per document condition and applied automatically
  • A confidence threshold and a review queue, with corrections logged
  • Daily inbound stabilised before the backfile is touched
  • One named owner for the queue, with time actually allocated to it

None of that requires a larger team. It requires the decisions to be made once, deliberately, and then enforced by the system rather than remembered by a person. That is the difference between an archive that has been scanned and an archive you can use.

Written byHelena CroftHead of Capture Services, Datacoll8

Book a demo

See it run against your own documents

Bring three or four representative agreements to the demo. We will run them through classification, extraction and verification live, and tell you plainly where the pipeline would need tuning for your paper.

Demos run Monday to Friday, 08:30 to 18:00 GMT

Expires in

Limited time offer

We rebuilt your site for you. Claim it and we handle everything transfer, hosting, and your domain. Then update it anytime, just by asking AI.

Host for only$8 per monthBilled yearly
Claim limited offer now