• Beyond the Spinner: What Should Analytical Agents Surface While They Work? - Vela
    📅 April 1, 2027✍️ Md. Rahat-Uz-Zaman, Srishti Palani, Dennis Bromley, Vidya Setlur📚 CHI 2027 (submitted)🎯 article
    Conversational and agent-driven analytics increasingly carry out long-running, multi-step workflows, yet users are often left with either opaque progress indicators or verbose execution traces. A formative study with data and analytics practitioners characterized what users want to see while an agent works, when they prefer to monitor versus intervene, and how these needs vary across tasks and contexts, yielding a framework for intermediate analytical surfacing. We present Vela, a system for adaptive intermediate surfacing that operationalizes this framework through a lightweight Observer that incrementally constructs grounded, task-specific analytical state and adapts disclosure based on task and interaction behavior. A summative study used Vela as a research probe to examine how people monitor, verify, and steer long-running analyses in practice. Together, these studies contribute an empirical characterization of intermediate visibility needs, design guidelines and a task-aware surfacing framework, and an observer-based architecture for adaptively constructing and disclosing intermediate analytical state.
  • Blabber: A Language for Chart Annotation - Blabber
    📅 April 1, 2027✍️ Huichen Will Wang, Md. Rahat-Uz-Zaman, Md Dilshadur Rahman, Andrew M. McNutt, Paul Rosen, Alper Sarikaya, Chenglong Wang📚 CHI 2027 (submitted)🎯 article
    Annotations help readers link patterns and evidence to an author's message. However, annotation authors often have to express intentions through verbose and fragile specifications that hard-code low-level parameters such as coordinates, lengths, and text offsets. We present Blabber, a declarative annotation language that allows users to express high-level annotation intent through operators and data references, leaving low-level configuration to a geometry-aware compiler. The Blabber compiler models annotation layout as a constraint optimization problem to identify configurations that balance clarity and simplicity. We implement Blabber as a library supporting 18 chart types and 13 annotation operators. We present a gallery of Blabber charts and case studies comparing Blabber with annotation features in major visualization libraries. On an LLM-driven chart annotation authoring benchmark, AnnoBench, GPT-5.6 Sol achieves a mean quality score of 11.59/15 with Blabber, compared with 7.00/15 when directly outputting Vega-Lite code and 5.81/15 when using AnnoGram, a baseline annotation language.
  • A Case Study in Accessible Redesign of a Wastewater Dashboard - Accessible Wastewater Dashboard
    📅 October 1, 2026✍️ Tingying He, Jake Wagoner, Md. Rahat-Uz-Zaman, Md Dilshadur Rahman, Willy Ray, George G. Vega Yon, Matthew Samore, Alexander Lex, Paul Rosen📚 IEEE VIS 2026 Workshop on Visualization for Accessibility🎯 article
    A case study of an accessible redesign of a public wastewater surveillance dashboard, presented at the IEEE VIS 2026 Workshop on Visualization for Accessibility.
  • Visual Accessibility Auditing Across Interaction States and Reading Orders - Site Accessibility Auditor
    📅 October 1, 2026✍️ Md. Rahat-Uz-Zaman, Jake Wagoner, Tingying He, Alexander Lex, Paul Rosen📚 IEEE VIS 2026 Workshop on Visualization for Accessibility🎯 article
    Automated accessibility checkers evaluate a single snapshot of a page, but many barriers emerge only over multiple interactions, not in a single static state. Several WCAG 2.2 criteria depend on geometry, focus state, or interaction history, such as a button that becomes obscured once a sticky header appears. Others arise when the visual layout, the screen-reader order, and the keyboard (Tab) order of a page disagree, so different users encounter the same content in conflicting sequences. Auditing these requires reasoning about geometry and structure alongside the live page, which makes it both a detection and a visualization problem. We present Site Accessibility Auditor, a Chrome DevTools extension with two coordinated lenses. The interaction lens automatically drives the page through its reachable states and overlays interaction-dependent issues as inspectable visual evidence on a full-page capture. The reading and focus order lens aligns the visual, screen-reader, and keyboard orders, shows where they diverge, and suppresses intentional patterns such as skip links. We evaluate the extension from two vantage points: as site owners and as an external auditor. In a design partnership with the Utah Wastewater Surveillance System dashboard, our tool motivated the team to ship accessibility fixes. We used the tool on the Our World in Data site, showing that the approach runs on an unmodified, in-the-wild site.
  • 3DMPE: 3D Multi-Perspective Embedding - 3DMPE
    📅 July 6, 2026✍️ Vahan Huroyan, Md. Rahat-Uz-Zaman, Stephen Kobourov📚 38th Canadian Conference on Computational Geometry (CCCG 2026)🎯 article
    We study 3D point cloud reconstruction from multiple partially observed 2D projections. Given two or more projections of an unknown 3D point cloud, together with cross-view point correspondences and visibility information, our goal is to recover a consistent 3D configuration when different views contain different subsets of points. We propose 3D Multi-Perspective Embedding (3DMPE), an optimization-based, training-free method that reconstructs the 3D point cloud and, in the variable-projection setting, jointly estimates the projection maps. 3DMPE extends Multi-Perspective Simultaneous Embedding to accommodate missing points and incomplete pairwise distance information across views. We consider both fixed-projection and variable-projection settings. Unlike learning-based reconstruction methods that infer shape from raw images and often depend on training data, 3DMPE operates on geometric observations with established correspondences and does not require category-specific training. Experiments on ShapeNet and Pix3D evaluate reconstruction quality using Chamfer Distance, Earth Mover Distance, and RMSE-Optimize-Align (ROA), and examine the effects of initialization, the number of views, point visibility, and several noise regimes, including noisy distances and erroneous correspondences. The results demonstrate that 3DMPE can effectively reconstruct point clouds from partial multi-view geometric observations.
  • ChannelExplorer: Exploring Class Separability Through Activation Channel Visualization - ChannelExplorer
    📅 July 1, 2026✍️ Md. Rahat-Uz-Zaman, Bei Wang, Paul Rosen📚 IEEE Transactions on Visualization and Computer Graphics, 32(7), 5442–5456🎯 article
    Deep neural networks (DNNs) achieve state-of-the-art performance in many vision tasks, yet understanding their internal behavior remains challenging—particularly how different layers and activation channels contribute to class separability. We introduce ChannelExplorer, an interactive visual analytics tool for analyzing image-based outputs across model layers, emphasizing data-driven insights over architecture analysis for exploring class separability. ChannelExplorer summarizes activations across layers and visualizes them using three primary coordinated views: a Scatterplot View to reveal inter- and intra-class confusion, a Jaccard Similarity View to quantify activation overlap, and a Heatmap View to inspect activation channel patterns. Our technique supports diverse model architectures, including CNNs, GANs, ResNet and Stable Diffusion models. We demonstrate the capabilities of ChannelExplorer through four use-case scenarios: (1) generating class hierarchy in ImageNet, (2) finding mislabeled images, (3) identifying activation channel contributions, and (4) locating latent states' position in Stable Diffusion model. Finally, we evaluate the tool with expert users.
  • AnnoBench: A Benchmark for Visualization Annotation Generation - AnnoBench
    📅 June 1, 2026✍️ Md. Rahat-Uz-Zaman, Md Dilshadur Rahman, Andrew McNutt, Paul Rosen📚 EuroVis 2026 (submitted)🎯 article
    Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for visualization annotation that materializes the inherent challenges of this domain in a structured and testable manner. AnnoBench pairs visualizations from professional data journalism and visualization galleries with annotation tasks, spanning four representation formats, five chart description conditions, and two prompt specification levels. The benchmark is executed via VLM-as-a-judge, using models aligned with manual human assessment. We evaluate the benchmark via four one-factor-at-a-time experiments, exploring the effects of input representation, semantic context, prompt specificity, and model selection on annotation quality. This work provides a foundation for advancing annotation automation, tooling, and visualization-generation pipelines.
  • AnnoGram: An Annotative Grammar of Graphics Extension - AnnoGram
    📅 August 8, 2025✍️ Md Dilshadur Rahman, Md. Rahat-Uz-Zaman, Andrew M McNutt, Paul Rosen📚 IEEE VIS 2025, pp. 236–240🎯 article
    Annotations are central to effective data communication, yet most visualization tools treat them as secondary constructs — manually defined, difficult to reuse, and loosely coupled to the underlying visualization grammar. We propose a declarative extension to Wilkinson’s Grammar of Graphics that reifies annotations as first-class design elements, enabling structured specification of annotation targets, types, and positioning strategies. To demonstrate the utility of our approach, we develop a prototype extension called Vega-Lite Annotation. Through comparison with eight existing tools, we show that our approach enhances expressiveness, reduces authoring effort, and enables portable, semantically integrated annotation workflows.
  • Prediction of Apple Leaf Diseases Using Multiclass Support Vector Machine
    📅 February 1, 2021✍️ Soarov Chakraborty, Shourav Paul, Md. Rahat-uz-Zaman📚 2021 2Nd international conference on robotics, electrical and signal processing techniques (ICREST)🎯 article
    Every year apple yield has been affected by Black rot and Cedar apple rust. It has a significant effect on both the apple industry and the country's economy. Here, we recommend a system to detect diseases from the infected apple leaves by combining machine learning and image processing principles. This approach can classify both infected and non-infected apple leaves efficiently. The identification is started by preprocessing the image using several image processing techniques, including the Otsu thresholding algorithm and histogram equalization. Using the image segmentation region of the infected part separates, and a Multiclass SVM recognizes the disease type from the original leaf image among 500 images with 96% accuracy. It also demonstrates the percentage of the total infected area of that diseased apple leaf image.
  • Visualizing Interaction Networks and Evidence in Biomedical Corpora - Blobviz
    📅 June 14, 2023✍️ Enrique Noriega-Atala, Md. Rahat-Uz-Zaman, Ruchika Bhat, Mladen Jergovic, Stephen G. Kobourov, Janko Nikolich-Zugich📚 Pacific (formerly Asia-Pacific APVIS) Visualization Symposium🎯 article
    The abundance of scientific articles published and indexed in publicly accessible repositories has spurred the research and development of automated information extraction systems. The output of such systems can be used to assemble large networks capturing the understanding of mechanistic pathways and their interactions as represented in the underlying body of research.We describe a system designed to help researchers search, visualize and interact with biological networks derived via information extraction tools. As input, the system takes a dataset of biological and biochemical interactions automatically generated by an information extraction system and provides an interface designed to search, visualize and interact with the data. The usage paradigm consists of identifying a starting point for a search, then using the data’s network structure by incrementally exploring the immediate neighborhood of the elements displayed by the system.Our system differs from prior work as it leverages both the network structure in the data and the natural language text backing those connections: every connection displayed is traceable back to the documents and phrases in the corpus that support that specific piece of information. We also present two case studies with immunobiology researchers using the system to find previously unknown relationships between biological entities. While the evidence suggesting these relationships already existed, it was scattered across the literature, and existing specialized web databases and domain-search engines could not find it. The system is open-source, with the code publicly available on GitHub.
  • A Multi-Modal Human Machine Interface for Controlling a Smart Wheelchair
    📅 December 14, 2019✍️ Saifuddin Mahmud, Xiangxu Lin, Jong-Hoon Kim, Hasib Iqbal, MD Rahat-Uz-Zaman, Sakib Reza, M Asifur Rahman📚 IEEE 7th Conference on Systems, Process and Control (ICSPC), 2019🎯 article
    As the number of disabled people all over the world is increasing very fast, the role of an electric wheelchair is becoming crucial to improve the mobility for them. Independent mobility is a vital aspect of self-respect and plays an important role in the life of a disabled person. The smart wheelchair is an endeavor to provide an self-supporting mobility to those people who are not able to move freely. Typical electric powered wheelchairs are usually controlled by the traditional joysticks which cannot fulfill the needs of a person who has motor disabilities and some specific types of disabilities like paralysis who can only move their eyes. This paper aims to develop a multi-modal human machine interface for the larger domain of disabled persons to control the wheelchair efficiently. The interface comprises joystick, smart hand-glove, head movement tracker and eye tracker. The system presented in this paper can support a wide variety of users with different types of disabilities.
  • Agglomerative Clustering of Handwritten Numerals to Determine Similarity of Different Languages
    📅 December 18, 2019✍️ Md. Rahat-uz-Zaman, Shadmaan Hye📚 22nd International Conference on Computer and Information Technology (ICCIT), 2019🎯 article
    Handwritten numerals of different languages have various characteristics. Similarities and dissimilarities of the languages can be measured by analyzing the extracted features of the numerals. Handwritten numeral datasets are available and accessible for many renowned languages of different regions. In this paper, several handwritten numeral datasets of different languages are collected. Then they are used to find the similarity among those written languages through determining and comparing the similitude of each handwritten numerals. This will help to find which languages have the same or adjacent parent language. Firstly, a similarity measure of two numeral images is constructed with a Siamese network. Secondly, the similarity of the numeral datasets is determined with the help of the Siamese network and a new random sample with replacement similarity averaging technique. Finally, an agglomerative clustering is done based on the similarities of each dataset. This clustering technique shows some very interesting properties of the datasets. The property focused in this paper is the regional resemblance of the datasets. By analyzing the clusters, it becomes easy to identify which languages are originated from similar regions.
  • Audio Future Block Prediction with Conditional Generative Adversarial Network
    📅 December 28, 2019✍️ Md. Rahat-uz-Zaman, Shadmaan Hye, Mahmudul Hasan📚 3rd International Conference on Electrical, Computer & Telecommunication Engineering, 2019🎯 article
    Signal processing is a vast subfield of electrical and computer science where audio signal processing has secured a remarkable position to restore corrupted or missing audio blocks. However, generating possible future audio block from the previous audio block is still a new idea that can help to reduce both audio noise and partially missing an audio segment. In this paper, a generative adversarial network (GAN) along with a pipeline is proposed for the prediction of possible audio after an input audio sequence. The proposed model uses short-time Fourier transformation of audio to make it an image. The image is then fed to a conditional GAN to predict the output image. After that, Inverse short-time Fourier transform is then applied to that predicted image, generating the predicted audio sequence. For a small audio sequence prediction, the proposed methodology is quite fast, robust and has achieved a loss of 0.43. So it is may work well if deployed on a voice call and broadcasting applications.
  • Extraction of Sequence from Bangla Handwritten Numerals and Recognition Using LSTM
    📅 June 7, 2020✍️ Shadmaan Hye, Md. Rahat-uz-Zaman, M. A. H. Akhand📚 IEEE Region 10 Symposium (TENSYMP), 5-7 June 2020, Dhaka, Bangladesh🎯 article
    In the promising era of Handwritten Numeral Recognition (HNR), despite Bangla being one of the major languages in the Indian subcontinent, fewer explorations have been done on Bangla numerals compared to other languages. Among the existing methods, several convolutional neural network (CNN) based method outperformed other methods. But CNN always gets confused with some specific Bangla numerals due to the similarity of shape and size of different numerals. The main purpose of this study is to expand Bangla HNR by considering a novel methodology with a Long Short-Term Memory (LSTM) network. In the proposed method, images are thinned and a sequence is extracted. These extracted sequences are used to classify using LSTM network. Both single-layer LSTM and Deep LSTM models are trained and performance tested on a benchmark dataset with a large number of samples. On the other hand, traditional CNN is also trained for better understanding. Experimental outcomes revealed that the proposed LSTM based method outperformed CNN with remarkable accuracy for the similar shaped numerals. Finally, the proposed method achieved a test set recognition rate of 98.03% which is better than or competitive to other prominent existing methods.
  • VISION-IT: A Novel Approach for Blind to Help hear what is in Front
    📅 ✍️ Md. Rahat-uz-Zaman, Shadmaan Hye📚 2020 24th International Conference on Computer and Information Technology (ICCIT)🎯 article
    Esc
    Search the whole portfolio Find publications, posts, projects, courses, and slides.