Find out if your data can support
what you want to build.

Most automation projects fail on the data, not the software. We test whether your documents can be read, whether your records are complete, and whether your systems will hand the data over. Then we score each source and tell you what to fix first.

This is for you if

Services

Every check we run on your data

All of this is measured on your own files. We do not ask your team to rate their own data, because everybody says it is fine until it is tested.

01

List every place the data actually lives

The ERP, the shared drive, the mailbox somebody has been forwarding to for six years, and the spreadsheet on one person's laptop. For each one we record who owns it, how much is in it, and how far back it goes.

02

Test how well your documents can be read

We take a real sample of your paperwork and run it through several document readers. We report accuracy field by field, not as one number, because a tool that reads the invoice total correctly every time may still miss the purchase order number half the time.

03

Count what is missing

Blank fields, duplicate records, records that contradict each other, and records that refer to something that no longer exists. We give you the count and the percentage, per source.

04

Check the same thing is named the same way everywhere

One supplier written four different ways, part codes with and without dashes, dates in two formats, weights in kilos in one system and pounds in another. This is the single most common reason a build takes longer than quoted.

05

Check you can get the data out

For each system, whether there is an API, a database view, a scheduled export, or nothing at all. If the only way out is a person clicking Export every morning, we need to know that now and not in month three.

06

Check how current the data is

How often each source updates, and how far behind it can be. Some processes need the data as it happens. Most are fine with once a day, and that difference changes what has to be built.

07

Check what you are allowed to do with it

Which fields count as personal data, whether they are allowed to leave your country, what has to be hidden before anything is sent to an outside service, and how long you are required to keep it.

08

Set aside a set of examples to test against later

We pick a spread of your real cases, including the awkward ones, and record the correct answer for each. Whatever gets built later is measured against this set. Without it, nobody can prove the automation works.

09

Score every source and say what to fix

Each source gets a rating, the work needed to lift it, and an honest note on whether that work is worth doing. Some sources are cheaper to replace than to clean.

Send us ten of your documents and we will tell you on a call how readable they are.

Our stack

The tools, standards and methods we use

You do not need to know any of these. They are listed so you can see we test data the way the industry does, with tools your own IT team will recognise.

Testing data qualityThese run the same checks every time and produce a pass or fail, so the result is a measurement rather than an opinion
Great ExpectationsSoda Coredbt testsCustom SQL checks
The six things we measureThe standard set used across the industry, so your scores mean the same thing to any supplier you show them to
CompletenessAccuracyConsistencyTimelinessValidityUniqueness
Reading documentsWe run your sample through more than one, because each is stronger on different layouts
Azure AI Document IntelligenceGoogle Document AIAWS TextractTesseract OCRField level accuracy scoring
Finding records that are really the same thingCatches the same supplier or part written four different ways, which is the most common hidden problem
Fuzzy matchingLevenshtein distanceSplinkDedupeGolden record
Standards for how data should be writtenAgreeing one format for dates, currencies and product codes now avoids a rewrite later
ISO 8601 datesISO 4217 currenciesGS1 and GTINUNSPSCHS codes
Getting the data out of your systemsWe check what already exists before assuming anything has to be built
RESTODataODBC and JDBCChange data captureDebeziumScheduled exports
Privacy and the rules that applySettled before any of your data is sent anywhere, not after
GDPRPersonal data classificationMasking and tokenisationData residencyRetention rules
Building the test setThe agreed set of correct answers that everything built later is measured against
Label StudioGold setHoldout setAgreement between reviewers

The report

Every source gets a score and a verdict

The centre of the report is one table. It tells you, source by source, whether you can build on it today, what it would take to make it usable, or whether to leave it alone.

SourceHow muchCan we read itCan we get it outVerdict
Supplier invoices, email inbox About 310 a month 94% of fields correct Yes, mailbox access Ready
Delivery notes, scanned at the gate About 900 a month 61% of fields correct Yes, shared drive Fix first
Supplier master list, ERP 2,400 records 18% are duplicates Yes, database view Clean up
Purchase orders, older system About 260 a month Structured, no reading needed Nightly export only Usable

Example figures. Your table will have your own sources and your own numbers.

Example

Why one source failed and what it cost to fix

Example

Take the delivery notes in the table above, the ones scoring 61%. On the face of it that looks like a document reading problem, so the obvious answer is a better tool. It was not.

When we looked at the failures, almost all of them came from one loading bay. The scanner there had been broken since March and the team had been photographing the notes on a phone instead, at an angle, often in poor light. The other bays scored above 90% on the same tool.

So the fix was not software. It was a scanner, and ten minutes telling the team to use it. That took the source from unusable to ready for about the price of the scanner, and it is the sort of thing you only find by looking at which documents failed rather than at the average.

This is an example, made up to show how the checks work. Your sources and numbers will be different.

FAQ

Questions we get asked before starting

How long does this take?

Usually about two weeks. Most of that is waiting for access to systems, so it goes faster if someone can arrange that before we start.

Do you need our live systems?

No. A sample of real documents and a read-only export is normally enough. We would rather not touch anything live at this stage.

Can we do this without sending you our data?

Yes. We can run the tests inside your own environment, or on a sample with the personal details removed. Some clients require this and it is a normal request.

What if the answer is that our data is not ready?

Then you have found out for the cost of a two week check rather than the cost of a failed build. The report tells you what to fix and in what order, and a lot of it is usually cheaper than people expect.

We already did this with another supplier. Is it worth repeating?

Only if their report told you accuracy per field and per source. Most do not. If yours gave one overall score with no breakdown, it will not tell you what to fix.

Do we have to build with you afterwards?

No. The report is yours, including the test set. You can hand it to your own team or to another supplier.

Next step

Send us ten documents and we will read them.

Pick ten of the documents your team handles most often, including two or three bad ones. We will run them and tell you on the call what came out correct and what did not.

Ahmedabad, India. We work with teams in the US, UK, Europe, Singapore and the Gulf, and we are used to the time difference.