Text Joins for Data Cleansing and Integration in an RDBMS

Luis Gravano
Panagiotis Ipeirotis
Nick Koudas
Divesh Srivastava

Venue: Proceedings of the 19th IEEE International Conference on Data Engineering (ICDE), 2003
Mar 2003
Status: Refereed
Type: Conference

An organization’s data records are often noisy because of transcription errors, incomplete information, lack of standard formats for textual data or combinations thereof. A fundamental task in a data cleaning system is matching textual attributes that refer to the same entity (e.g., organization name or address). This matching can be effectively performed via the cosine similarity metric from the information retrieval field. For robustness and scalability, these “text joins” are best done inside an RDBMS, which is where the data is likely to reside. Unfortunately, computing an exact answer to a text join can be expensive. In this paper, we propose an approximate, sampling based text join execution strategy that can be robustly executed in a standard, unmodified RDBMS.

Data Cleaning

Panos Ipeirotis

Text Joins for Data Cleansing and Integration in an RDBMS

Panos Ipeirotis

Text Joins for Data Cleansing and Integration in an RDBMS

Related Files:

Related Projects: