Report ID
2003-01
Report Authors
S. Alireza Aghili, Divyakant Agrawal, Amr El Abbadi
Report Date
Abstract
The problem of proximity search in biological databases is addressed. We study vector transformations and conductthe application of DFT(Discrete Fourier Transformation) and DWT(Discrete Wavelet Transformation, Haar) dimensionalityreduction techniques for DNA sequence proximity search to reduce the search time of range queries. Our empiricalresults on a number of Prokaryote and Eukaryote DNA contig databases demonstrate up to 50-fold filtration ratio of thesearch space, and up to 13 times faster filtration. The proposed transformation techniques may easily be integrated asa preprocessing phase on top of the current existing similarity search heuristics such as BLAST, PattenHunter, FastATA,QUASAR and to efficiently prune non-relevant sequences. We study the precision of applying dimensionality reductiontechniques for faster and more efficient range query searches,and discuss the imposed trade-offs.
Document
2003-01.pdf237.65 KB