Supervised learning based instance matching system for data linking
作者
Gulshakh Kaur,Poonam Saini,Shilpa Verma
标识
DOI:10.1109/icacci.2017.8125813
摘要
In the Linked Data context, identity link is one of the most important semantic links that can be established between the datasets. It specifies that different identifiers refer to the same real world object and therefore must be linked. The process of detecting these identical instances across different data repositories is referred as instance matching. This is used to connect existing data sources and provides effective data integration from multiple data sources, therefore, maintains consistency and integrity of resultant data. To establish links, an instance matching system follows a link configuration which specifies the properties, similarity measures and other parameters required for data linking. When the data is huge detecting the configuration manually is not feasible. The paper proposes the supervised learning based instance matching system that relies on the learning of link configuration to establish identity links. The output of the learning algorithm is the optimal link configuration which returns the best possible combination of linking parameters.