Skip to content
Printed financial statement with one amount circled in red and two marker pens beside it
Knowledge base

Data quality: measure, improve and monitor

Tom Frohn Co-founder of Optilise
5 min read

Data quality is the degree to which your data is fit for the purpose you use it for. You assess it on six dimensions: accuracy, completeness, timeliness, consistency, uniqueness and validity. Improving it comes down to four moves: decide which data is critical, measure how good it is now, fix the errors in the source system and check every month that it stays that way.

“Is this figure actually right?” Hear that question once in a meeting and you will hear it again with every report. And each time, the answer costs an afternoon of digging instead of a decision.

Then the problem is rarely the dashboard, it is what sits underneath. A report built on unreliable data does not get used a little less, it stops being used entirely from the first error someone finds. That is why every BI project that lasts starts with the quality of the data, not with the design of the screen.

Below you will find what data quality is and how to measure and improve it, plus two things most explanations leave out: why you fix errors at the source rather than in Power BI, and how to make data quality visible in the dashboard itself.

What is data quality?

Data quality is not a fixed property of data, but a relationship between the data and what you want to do with it. For revenue by region it hardly matters whether a phone number is missing, for a calling campaign it is everything. The same customer file is excellent for one question and useless for another.

So the first question is never “is our data good?”, but “good enough for what?”. Skip that question and you start a clean-up project with no end, because there is always another field that could be better.

Data quality is also different from data governance. Data quality describes the state of the data: is it correct, complete, current. Governance is the set of agreements that protect that state: who owns it, which definition applies, who may change what. Without those agreements, quality slides back after every clean-up. Both are part of data management, the full set of agreements for handling data from entry to deletion. This article goes deeper into one part of it.

The six dimensions of data quality

To make “good enough” measurable, data quality is split into dimensions. Every framework uses a slightly different list, but these six come back everywhere:

  • Accuracy. Does the data match reality? A customer who has moved but is still listed at the old address is complete and still wrong.
  • Completeness. Are the fields you need filled in? An order without a cost centre ends up in no departmental report.
  • Timeliness. Is the data available in time and still current? Last night’s stock level is fine for a weekly report, not for the salesperson promising a delivery date today.
  • Consistency. Is the same data identical in every system? If a debtor has a different customer number in the accounts than in the CRM, your reports will not reconcile.
  • Uniqueness. Does each record appear once? Duplicate customers split revenue across two lines and make every top ten unreliable.
  • Validity. Does the data meet the agreed format and rules? A postcode with too many digits or an order date in 2062 is filled in, but not valid.

Some frameworks add context as a seventh: a revenue figure can pass all six and still mislead if nobody knows whether discounts have been deducted. Definitions like that belong with your master data.

What does poor data quality cost?

Woman at a desk with her hands in her hair, next to stacks of files and papers

Little per error and a lot in total. The cost never shows up as one line in the budget, it is spread across hours, decisions and missed revenue. Which is exactly why nobody adds it up.

  • Manual work. The controller who spends a day every month correcting exports before the report can go to the management team. Twelve working days a year on work that good input would have made unnecessary.
  • Wrong decisions. A product that looks loss-making because the purchase price was never updated gets dropped from the range. The error sits in one field, the decision hits the whole product line.
  • Missed revenue. Customers left out of a campaign because their industry field is empty.
  • Legal duty. Under the GDPR you have to keep personal data accurate and update it where necessary, and people can ask you to correct it, as the Dutch Data Protection Authority explains.
  • Trust. The most expensive item. One figure that is shown to be wrong, and people go back to their own spreadsheets. Now you are paying for two versions of the truth.

Just for fun, search your CRM for a customer called “Test”. In most systems it has been there since the implementation, at 1 Test Street, with a few trial orders that have been quietly counted in revenue for years.

How do you measure data quality?

Woman wearing glasses looks at a spreadsheet on her laptop and takes notes in a notebook

By writing a rule you can count for each critical data element. Not “the customer data is messy”, but “312 of the 4,180 active customers have no company registration number”. Only with a number like that can you see whether things improve.

Start with the data your most important decisions rely on, for most mid-sized companies customers, products and orders. Write one measurement rule per dimension:

DimensionExample of a measurement rule
AccuracyNumber of parcels returned because of a wrong address, per month
CompletenessPercentage of active customers with an industry filled in
TimelinessHours between the last change in the source and the last refresh of the report
ConsistencyNumber of debtors in the accounts with no match in the CRM
UniquenessNumber of customers sharing a registration number or email address
ValidityNumber of orders with a delivery date before the order date

If you work with Power BI, you already have a first measuring tool. Data profiling in Power Query shows per column which share is valid, an error or empty. By default it only looks at the first 1,000 rows, so switch the profiling to the entire data set.

That first measurement is your baseline. Record it with a date, or every later discussion about progress becomes one gut feeling against another.

Improving data quality in five steps

Two people at a table filling in an online form together on a laptop

Cleaning up is the easy part, stopping the mess from coming back is the work. That is why the plan does not end at the clean-up.

  1. Define the purpose and the critical data. Which decision needs to improve, and which fields do you need for it?
  2. Measure where you stand. Most errors usually come from two or three causes.
  3. Clean up once, in the source system. Merge duplicate customers, fill empty mandatory fields and correct invalid values. Do it in the system where the data originates, not in a copy.
  4. Control input at the front door. Mandatory fields, drop-down lists instead of free text, a check on the company registration number when a customer is created. Every error you stop at input is one you never have to hunt for later.
  5. Assign ownership. One named person per data domain, not a department. That owner decides what a valid value is and records which system leads when two sources disagree.

How these agreements fit into a wider plan, with ownership per source and per definition, is covered in our article on data strategy.

Fix it at the source, not in Power BI

The most tempting answer to bad data is the fastest one: fix it in the report. A replace step in Power Query that turns “Ltd” and “Limited” into one, a correction table for three products with the wrong cost price. The dashboard adds up again and everyone is happy.

Except that every correction in the report hides the problem in exactly the place where you should have seen it. The source stays wrong, and so do the invoice and the export to the accountant. Meanwhile the correction steps pile up. After two years there are forty, nobody remembers why step seventeen exists and nobody dares to delete any of them.

Combining, converting and modelling data does belong in Power BI, or in a data warehouse underneath it. The distinction: a transformation changes the shape of correct data, a correction changes the content of wrong data. The first belongs in the reporting layer, the second in the source system.

Making data quality visible in your dashboard

What nobody sees, nobody fixes. That is why a simple measure often works best: a separate report page with the measurement rules from above, as an exception list.

Such a page shows a score per data domain and the trend since the baseline, with the records that fail below it and the owner’s name on every line. That turns data quality from an abstract problem into a work list someone can get through on a Monday morning.

Two conditions: the list only contains errors someone can fix in the source, otherwise it becomes a wall of complaints, and it gets fifteen minutes on the agenda every month in the meeting where the figures are discussed. How to build a page like that into a report the management team steers on is covered in our article on getting a Power BI dashboard built.

When not to make it a separate project

Not every organisation needs a data quality project. If you work with one system, a small team enters the data and there is never a debate about which figure is right, a fixed half-hour check each month is enough. An outside party or a tool would cost you more than it returns.

That changes once two systems hold the same data, reports are corrected by hand before they go to the management team, or someone postpones a decision because the figure cannot be trusted. Then the question is no longer whether you tackle it, but how long you keep accepting the cost of leaving it.

So how does it work for you: are errors in your figures fixed in the system where they arise, or again every month in the report? If you want to know where your first leak is, get in touch.

Frequently asked questions

What is data quality?

Data quality is the degree to which data is fit for the purpose you use it for. You assess it on dimensions such as accuracy, completeness, timeliness, consistency, uniqueness and validity. The same file can be good enough for one question and useless for another.

What are the dimensions of data quality?

The six that appear in almost every framework: accuracy (does it match reality), completeness (are the fields you need filled in), timeliness (is it available in time and still current), consistency (is it the same in every system), uniqueness (does each record appear once) and validity (does it meet the agreed format). Some frameworks add context as a seventh.

How do you measure data quality?

By writing a countable rule for each critical data element, such as the number of active customers without an industry or the number of orders with a delivery date before the order date. The first measurement is your baseline. In Power BI, Power Query's data profiling shows the share of valid, error and empty values per column, provided you set it to the entire data set.

How can you improve data quality?

Decide which data is critical for your most important decisions, measure where it stands now, clean it once in the source system, control input with mandatory fields and drop-down lists, and give each data domain an owner. After that, check every month that it stays that way.

What is the difference between data quality and data governance?

Data quality describes the state of the data: is it correct, complete and current. Data governance is the set of agreements that protect that state: who owns it, which definition applies and who may change what. Without governance, quality slides back after every clean-up.

What does poor data quality cost?

Rarely a single line in the budget, which is why nobody adds it up. The cost sits in manual work to fix exports, in decisions made on wrong figures, in missed revenue and in trust in your reports. Once one figure is shown to be wrong, people go back to keeping their own spreadsheets.

Who is responsible for data quality?

The business, not IT. Each data domain, such as customers or products, should have one named owner. That person decides what a valid value is and which system leads when two sources disagree. IT provides the checks and the integrations.

Why does data quality matter for AI?

An AI model learns from the data you give it and repeats the errors in it, only faster and with more confidence. Duplicate customers, outdated prices or empty fields lead to predictions that look convincing and are still wrong. Data quality is a precondition for AI, not a follow-up step.

Back to the knowledge base Back to top

Do you want to grow with data?

Curious about what we can do for you? We are happy to show you how data can help your organization grow.

Tom Frohn, Optilise
Get in touch