Posted Tuesday April 23, 2019

Bad Data

New AI tool able to better identify bad data

Researchers have developed a novel tool for managing the quality of your data.

The revolutionary tool, HoloClean is the first to use artificial intelligence (AI) to sift out dirty data and correct errors before processing it.

“More and more machines are making decisions for us, so all our lives are touched by dirty data daily,” said Ihab Ilyas, a professor in Waterloo’s David R. Cheriton School of Computer Science. “If organizations like banks or utility companies are working with bad data, it could negatively impact things such as credit scores or mortgage approvals.”

The HoloClean system, developed by Ilyas and colleagues from the University of Wisconsin and Stanford University, successfully tackles the problem of insufficient training data by automatically generating bad examples without messing with the information, but enough to train the system to find errors and correct them on its own.

After the AI is trained, it can then independently figure out what’s an error, what’s not, and if there’s an error determine the most probable value for the missing data. Users will then have a cleaner dataset to use in their analytics which will produce more trustworthy results.

“This work deviates from the old way of manually trying to clean the data which was expensive, didn’t scale, and does not meet the current needs for cleaning the data,” said Ilyas of Waterloo’s Faculty of Mathematics. “This system addresses the problem where the information is out there, and people are using it to run analytics, but it is not correct. It doesn’t provide information that was not there, but instead corrects information you assume is correct.”

The next step for the researchers is to pair error detection and data repair in one end-to-end solution for the ultimate data quality dashboard.

The paper detailing the error detection module titled “HoloDetect: Few-shot Learning for Error Detection”, which is slated to appear in Proceedings of the 2019 ACM SIGMOD conference is authored by Ilyas, Waterloo’s PhD candidate in Computer Science, Alireza Heidari, assistant professor in the Department of Computer Sciences at the University of Wisconsin-Madison, Theodoros Rekatsinas, and academic researcher at the University of Wisconsin, Joshua McGrath. is distributed twice weekly; Tuesday and Thursday

Not a Subscriber? Receive more articles on business issues
Sent directly to your email in-box.

contact us

Publisher is Exchange Business Communication Inc.
No part of this article may be reproduced without written permission from the publisher.
Email Publisher

ISSN 0824-45
Copyright, 2019

Share Your Story: pressrelease @
Expand Your Reach:
advertise @
Publisher: Exchange Business Communication Inc., PO Box 248, Waterloo, Ontario, Canada
Email Editor

Content published on this site represents the opinion of the individual, organization and/or source provider. is a online daily journal of Exchange Business Communication Inc. (1997) Publishers of Exchange Magazine est. 1983. Privacy Policy. Copyright of Exchange produced editorial is the copyright of Exchange Business Communications Inc. 2019. Submitted editorials, comments and releases are copyright of respective source(s).