Welcome to D
SIGMOD 2005
PODS 2005
SIGMOD-RECOR
CIDR 2005
CIKM 2005
COMAD 2005
CVDB 2005
DaMoN 2005
Data Enginee
DEBS05
DMSN 2005
DOLAP 2005
GIR 2005
GIS 2005
Hypertext 20
ICDE 2005
ICDM 2005
IHIS 2005
IQIS 2005
JCDL 2005
KRAS 2005
MDM 2005
MIR 2005
MobiDE 2005
P2PIR 2005
RIDE 2005
SBBD 2005
SIGIR 2005
SIGIR-FORUM
SIGKDD 2005
SIGKDD-EXP
<<< = SIGKDD-EXP P>>>
SSDBM 2005
TIME 2005
TKDE 2005
TODS 2005
VLDB 2005
VLDBJ 2005
WebDB 2005
WIDM 2005

Introduction: Special Issue on Link Mining


Lise Getoor and Christopher P. Dieh

  View Paper (PDF)  

Return to December 2005, Volume 7, Issue 2


Abstract

An emerging challenge for data mining is the problem of mining richly structured datasets, where the ob jects are linked in some way. Many real-world datasets describe a variety of entity types linked via multiple types of relations. These links provide additional context that can be helpful for many data mining tasks. Yet multi-relational data vio- lates the traditional assumption of independent, identically distributed data instances that provides the basis for many statistical machine learning algorithms. Therefore, new ap- proaches are needed that can exploit the dependencies across the attribute and link structure. With linked data, new challenges emerge in traditional data mining tasks and new inference tasks are introduced. Clas- sification of data ob jects, for example, is more complex with linked data, since the classification decisions are now depen- dent. This leads to the new challenge of collective classifica- tion, in which the classification decisions are resolved jointly. The introduction of links implies a host of new descriptive and predictive inference tasks based on the link structure. These range from capturing descriptive patterns in the link structure and predicting emerging links to discovery of clus- ters or communities based on attributes and link structure. In recent years, there has been a surge of interest in this area, fueled largely by developments on the Internet. The Internet has provided a medium for hundreds of millions of people to connect with others, communicate ideas, express value judgements through the organization of those ideas, and engage in commerce. This has led to significant inter- est in challenges such as mining the web for information retrieval, online social networks for fostering relationships, and transaction patterns for recommending items of poten- tial interest. Rapid advances in sensing, computation, and communication have spawned increasing volumes of multi- relational data. In domains such as biosurveillance, law enforcement, homeland security, and e-commerce, there is a continuing need for algorithms that can identify notable patterns in real-time transactional and event data. We ex- pect that the demands for algorithms and processes to help understand and make predictions based on multi-relational data will only increase.


©2006 Association for Computing Machinery