BDSLAB
University of Houston
Big Data Systems for AI Lab
Department of Computer Science

What we work on

Research

Selected ongoing and past projects in the group. Broad themes include monitoring neural network, optimizing neural network, sparse matrix algorithms using database-style coordinate storage, big data analysis, scalable algorithms, medical data mining, and information retrieval in database systems.

I/O-Efficient Algorithms for Sparse Matrix Multiplication and Addition

Hashirul Quadir

Matrix multiplication and addition sit at the core of data science, yet most libraries optimize for dense matrices in C behind NumPy and PyTorch. Real workloads — text, graphs, binary features — are massively sparse, and the usual CSR/CSC formats that work well in RAM and cache perform poorly for I/O. We show that external-memory algorithms inspired by database query processing balance CPU and I/O effectively and work well with GPUs. Experiments confirm that large sparse matrices need specialized algorithms, and that database-style approaches are competitive with traditional HPC methods.

Inspecting Neural Networks with Queries

Dipta Chandra Paul · Hashirul Quadir · Nicholas Anderson

Diagnosing a training neural network is usually limited to watching loss curves, while deeper inspection relies on custom scripts that are hard to interpret and costly to run. We build a prototype that instead monitors a network through SQL queries over a relational database, storing only binary neuron activity, biases and predictions — never the weight matrices — to keep high-frequency parameter transfer cheap. Analysts can then write standard SQL to study the decision boundary, the distribution of biases across hidden layers, and convergence behavior during training.

Sparse Matrix Algorithms for Evolving Neural Networks

Carlos Ordonez

Massive neural networks and constantly changing data sets are pushing matrix computation beyond what main-memory-centric HPC techniques handle well. We survey three key problems: keeping a large data set updated under frequent matrix entry insertions and deletions, sparse matrix addition and multiplication, and recomputing a deep network when a sparse data set changes. We propose parallel, I/O-efficient algorithms that store and process matrices as coordinate tuples — like a relational table — and argue this representation can complement, and potentially replace, dense arrays and compressed row/column formats for updating, explaining and monitoring evolving networks.

Scalable Data Summarization to Compute Machine Learning Models

Sikder Tahsin Al-Amin

We generalize a data summarization matrix to produce one or multiple summaries that benefit a broad class of models. It runs in R and Python on a single machine and in parallel on shared-nothing architectures, evaluating the demanding vector–vector multiplication in C++ behind a simple high-level function call — faster and simpler than Spark (MLlib) or a parallel DBMS computing the same models.

Graph Analysis on Distributed Database Systems

Xiantian Zhou

Graph analytics is among the most compute-intensive big data tasks. We design and optimize graph algorithms on distributed database systems — computing metrics such as betweenness centrality, transitive closure and triangle counting — and build scalable graph algorithms from scratch in C++.

Continuous Network Monitoring

Quangtri Thai

Testbeds built on platforms like Raspberry Pi lack flexible data-analysis capability. We model network instrumentation as streaming data processing, express it in SQL, and implement a lightweight database- and SQL-based monitoring system on edge devices in a distributed architecture.

Interactive Compiling of C/C++ in Python

Quangtri Thai

To combine Python's simplicity with the speed of C/C++, we wrap compute-intensive code in C/C++ using SWIG for seamless use from Python, keeping the wrapped code easy to maintain and callable with Python's intuitive syntax.