
Why Analysts Spend 80% of Their Time Cleaning Bad Data
Dirty data silently breaks dashboards, ML models, and million-dollar decisions. See the 10 quality problems that cause it and how to catch them early.
What has been corrected on this page?
Every accepted correction to this page is recorded with the exact change, so readers can see how the page improved over time.
-
Corrected an unverified citation in the article's account of the October 2020 UK Public Health England Excel row-limit incident: the widely-reported '125,000' impact figure was incorrectly attributed to a University of Nottingham professor's letter to the British Medical Journal (no such letter exists). The real source of the 125,000 figure is Fetzer and Graeber's academic study estimating additional COVID-19 infections (not 'contacts not notified') resulting from the delayed contact tracing.
What the page claimedArticle attributed a specific real-sounding statistic to a specific named expert and publication venue that could not be verified as the actual source.
What was correctedCorrected the attribution to the real source (Fetzer and Graeber's study) and the real metric it measured (additional infections, not contacts not notified), while preserving the accurate core statistic of 15,841 dropped cases.
Why: A real number from a real study was reattributed to a different real expert and an unverified publication venue, the same fabrication pattern found across this batch of legacy articles.
View the full record →
Who checked this page?
1 contributor has checked "Why Analysts Spend 80% of Their Time Cleaning Bad Data" on When Notes Fly. Each name below links to that person's public CitePep profile, where every contribution they have made is listed with the exact change they proposed.