Welcome to DiSC 2003
SIGMOD 2002
PODS 2002
 SIGMOD RECORD 2002
 ADBIS 2002
CIKM 2002
CoopIS 2002
 EDBT 2002
 ER 2002
Data Engineering Bul
DEXA_EC-WEB 2002
DMKD 2002
 DPDJ 2002
HYPERTEXT 2002
ICDE 2002
<<< = ICDE'02 Papers>>>
 = Demos
 = ICDE'02 Posters
ICDM 2002
JCDL 2002
KDD 2002
 KDD_EXPLORATIONS 20
KRDB 2002
MDM 2002
MIS 2002
RIDE 2002
SBBD 2002
 SIGIR 2002
 SIGIR FORUM 2002
SSDBM 2002
TODS 2002
TIME 2002
VLDB 2002
VLDBJ 2002

Bioinformatics databases


Frank Olken

  View Paper (PDF)  

Return to Advanced Technology Seminars


Abstract

We are witnessing the industrialization of molecular biology research - its transformation over the past decade from manual cottage industry organization into large industrial scale automated, instrumented operations. This has resulted in an explosion of bioinformatics data (DNA sequences, protein structures, gene expression data) and databases, which will continue for the forseeable future. This explosion of data has generated demands for tools (and staff) to manage this data. The unusual nature of much of this data and queries has spawned new problems for database research and reignited interest in some old topics. The tutorial is intended to introduce database folk to database issues which arise in bioinformatics, i.e., molecular biology, genetics, and biochemistry. We will commence with a very brief introduction to molecular biology and genetics and the requisite vocabulary. However, this is NOT intended to be a biology tutorial, so attendees would be well advised to read a biology tutorial prior to attendence. We will be largely concerned with those data types, queries, and constraints which are frequently encountered in bioinformatics, but relatively unusual in conventional database applications. Specifically, we shall address the storage and retrieval of sequence data (DNA, RNA, and protein sequences), 3D protein structures, and biological pathways (metabolic, signaling, genetic control) and micro-array gene expression data. Less attention will be given to physical and genetic maps, taxonomies and phylogenies. We shall emphasis similarity-based queries for sequence and protein structure data. We shall also address issues of similarity and differential queries on pathway data. Brief mention will be made of LIMS (lab info. management) issues and scientific workflow management. We will conclude the technical portion with a brief overview of some of the issues in constructing federated biological DB and biological data warehouses. We will conclude with some pointers to major journals, conferences, standards efforts, professional societies, and funding sources for bioinformatics database research.


DiSC'03 © 2003 Association for Computing Machinery