About
About me
I am a PhD student in Computer Science at the University of Houston and a Teaching Assistant in the Department of Computer Science, working with the Big Data Systems for AI Lab under Prof. Carlos Ordonez. My work sits at the intersection of machine learning, data science and parallel and distributed computing — building database-inspired systems to train, inspect and accelerate neural networks on large data sets. Currently, I am extending this work to monitor recurrent (LSTM) and Transformer networks.
Before joining UH, I taught as a Lecturer in the Department of CSE at Leading University, Bangladesh, and worked for a few years as a software engineer at WellDev.
Interests
Research interests
- Primary — machine learning, data science, artificial intelligence, parallel and distributed computing.
- Secondary — optimization, software engineering, algorithms, reinforcement learning.
Selected work
Publications
Inspecting Neural Networks with Queries
Dipta Chandra Paul, Hashirul Quadir, Nicholas Anderson, Carlos Ordonez · DEXA 2026
A lightweight relational-database framework that inspects neural networks through SQL queries, tracking neuron activations and biases to detect misclassifications with minimal time overhead.
A Generic Algorithm Integrating Parallel Processing and Block-based I/O for Large-Scale AI Analytics
Dipta Chandra Paul, Hashirul Quadir, Ladjel Bellatreche, Carlos Ordonez · under review, Data & Knowledge Engineering (DKE)
A parallel algorithm that combines block-based I/O with data partitioning to accelerate AI model execution in Python, enabling processing of very large data sets beyond main-memory limits.
Coursework
Course projects
NeuroDB: Monitoring Neural Networks as Data Evolves
COSC 6340 — Database Theory
NeuroDB is a database-oriented system for storing, tracking and analyzing how neural network parameters change as the underlying data updates over time. Networks are trained externally and the database is used strictly for monitoring and analysis: each model version is stored as relational data, parameter-level deltas are computed between successive versions, and SQL queries identify unstable layers and neurons. By treating time and versioning as explicit data, NeuroDB brings transparency, auditability and query-driven insight into how a model evolves — including the case where a data set's input dimensions periodically change by a little.
Urban AQI Forecasting in Houston: An End-to-End Pipeline and Comparative Modeling Analysis
COSC 6339 — Big Data Analytics
An end-to-end big data pipeline for collecting, preprocessing and exploring multiple data sources to forecast the US Air Quality Index in Houston, built to operate under strict memory constraints. The pipeline spans five stages — collection, preprocessing, exploration, training and validation — and compares sequence models (LSTM) against strong tabular baselines. The results suggest deep sequence models are a plausible direction for AQI forecasting, though in this implementation tree-based methods generalized better on unseen data.
Model-Based Text Compression Using LLM Next-Word Prediction
COSC 6397 — Topics Computer Science
Large language models are strong sequence predictors, which makes them promising for lossless text compression. While recent methods like LLMZip pair LLM predictions with arithmetic coding, other entropy coders are less explored. This project studies how LLM-based probability predictions combine with Huffman coding, ANS and arithmetic encoding, comparing compression ratio, bit rate and computational efficiency against non-entropy methods such as zstd. The results show that zstd, though simpler, reaches the lowest bit rate given the skewed nature of LLM-predicted distributions, while LLM inference remains the computational bottleneck — highlighting the efficiency–latency trade-offs and pointing to LLM compression as promising for archival storage, with smaller custom models needed for general use.
Teaching
Teaching
Teaching Assistant · University of Houston
- COSC 3380 — Database Systems, Fall 2025, with Prof. Carlos Ordonez.
- COSC 3380 — Database Systems, Spring 2026 & Summer 2026, with Dr. Victoria Hilford.
Before UH · Leading University, Bangladesh
- Lecturer, Department of CSE (2022–2025).
- Coach, competitive programming teams (2018–2022).
Education
Education
- PhD in Computer Science — University of Houston, USA (2025–2030, in progress).
- BSc in Computer Science and Engineering — Leading University, Sylhet, Bangladesh (2014–2018).
Toolbox
Technical skills
- Programming — C++, Python, Java, JavaScript, MATLAB.
- Machine learning — scikit-learn, TensorFlow, PyTorch, Pandas, NumPy, GeoPandas.
- Tools & frameworks — Git, Docker, Linux, LaTeX, Angular, React Native, Node.js.
Get in touch
Contact
Email: dpaul6@cougarnet.uh.edu