![]() ![]() ![]() |
![]() |
|
![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]() ![]()
|
Return to KEYNOTES Integration of data from multiple sources is one of the longest standing problems facing the database research community. In addition to being a problem in most enterprises and in large-scale science projects, research on this topic has been fueled by the promise of querying the WWW. I will begin by highlighting some of the significant recent achievements in the field of data integration, and will then focus on what I consider to be the main challenges going forward, namely, large-scale reconciliation of semantic heterogeneity, and on-the-fly information integration. To address these challenges, I argue for an approach based on computing statistics over corpora of database structures. At a fundamental level, the key challenge in data integration is to reconcile the semantics of disparate data sets, each expressed with different database structures. Computing statistics over a large corpus of schemas and mappings offers a powerful methodology for producing semantic mappings, the expressions that specify such reconciliation. In essence, the statistics offer hints about the semantics of the symbols in the structures, thereby enabling to detect when two symbols, from disparate schemas, should be matched to each other. The same methodology can be applied to several other data management tasks that involve search in a space of complex structures. I will illustrate several examples where this approach has been successful. ![]() ©2005 Association for Computing Machinery |