![]() ![]() ![]() |
![]() |
|
|
![]() ![]() ![]() ![]() ![]() |
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() |
Return to December 2005, Volume 7, Issue 2 An emerging challenge for data mining is the problem of mining richly structured datasets, where the ob jects are linked in some way. Many real-world datasets describe a variety of entity types linked via multiple types of relations. These links provide additional context that can be helpful for many data mining tasks. Yet multi-relational data vio- lates the traditional assumption of independent, identically distributed data instances that provides the basis for many statistical machine learning algorithms. Therefore, new ap- proaches are needed that can exploit the dependencies across the attribute and link structure. With linked data, new challenges emerge in traditional data mining tasks and new inference tasks are introduced. Clas- sification of data ob jects, for example, is more complex with linked data, since the classification decisions are now depen- dent. This leads to the new challenge of collective classifica- tion, in which the classification decisions are resolved jointly. The introduction of links implies a host of new descriptive and predictive inference tasks based on the link structure. These range from capturing descriptive patterns in the link structure and predicting emerging links to discovery of clus- ters or communities based on attributes and link structure. In recent years, there has been a surge of interest in this area, fueled largely by developments on the Internet. The Internet has provided a medium for hundreds of millions of people to connect with others, communicate ideas, express value judgements through the organization of those ideas, and engage in commerce. This has led to significant inter- est in challenges such as mining the web for information retrieval, online social networks for fostering relationships, and transaction patterns for recommending items of poten- tial interest. Rapid advances in sensing, computation, and communication have spawned increasing volumes of multi- relational data. In domains such as biosurveillance, law enforcement, homeland security, and e-commerce, there is a continuing need for algorithms that can identify notable patterns in real-time transactional and event data. We ex- pect that the demands for algorithms and processes to help understand and make predictions based on multi-relational data will only increase. ![]() ©2006 Association for Computing Machinery |