I have published and presented my work at prestigious conferences as well as workshops.

You can click on the keywords below to explore my research in specific areas.








View all publications in Google Scholar. Only selected publications are shown on this page currently. Click to


📰 Conference Papers


  1. SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
    Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Anik, Munem Shahriar, Mohsin Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez
    ICLR 2026 (A*) â–Ș [PDF]

    Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision–language models (VLMs). Prior work largely relied on synthetic or LLM-generated environments with limited task designs and puzzle-like setups, failing to capture the real-world complexity, visual noise, and diverse spatial relationships that VLMs encounter. To address this, we introduce SpatiaLab, a comprehensive benchmark for evaluating VLMs’ spatial reasoning in realistic, unconstrained contexts. SpatiaLab comprises 1,400 visual question–answer pairs across six major categories: Relative Positioning, Depth & Occlusion, Orientation, Size & Scale, Spatial Navigation, and 3D Geometry, each with five subcategories, yielding 30 distinct task types. Each subcategory contains at least 25 questions, and each main category includes at least 200 questions, supporting both multiple-choice and open-ended evaluation. Experiments across diverse state-of-the-art VLMs, including open- and closed-source models, reasoning-focused, and specialized spatial reasoning models, reveal a substantial gap in spatial reasoning capabilities compared with humans. In the multiple-choice setup, InternVL3.5-72B achieves 54.93% accuracy versus 87.57% for humans. In the open-ended setting, all models show a performance drop of around 10–25%, with GPT-5-mini scoring highest at 40.93% versus 64.93% for humans. These results highlight key limitations in handling complex spatial relationships, depth perception, navigation, and 3D geometry. By providing a diverse, real-world evaluation framework, SpatiaLab exposes critical challenges and opportunities for advancing VLMs’ spatial reasoning, offering a benchmark to guide future research toward robust, human-aligned spatial understanding. SpatiaLab is available at: https://spatialab-reasoning.github.io/.
    @inproceedings{wasi2026spatialab,
    title={SpatiaLab: Can Vision{\textendash}Language Models Perform Spatial Reasoning in the Wild?},
    author={Azmine Toushik Wasi and Wahid Faisal and Abdur Rahman and Mahfuz Ahmed Anik and Munem Shahriar and Mohsin Mahmud Topu and Sadia Tasnim Meem and Rahatun Nesa Priti and Sabrina Afroz Mitu and Md. Iqramul Hoque and Shahriyar Zaman Ridoy and Mohammed Eunus Ali and Majd Hawasly and Mohammad Raza and Md Rizwan Parvez},
    booktitle={The Fourteenth International Conference on Learning Representations},
    year={2026},
    url={https://openreview.net/forum?id=fWWUPOb0CT}
    }
  2. Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
    Azmine Toushik Wasi, Mst Rafia Islam, Mahfuz Ahmed Anik, Md Manjurul Ahsan, TH Rafi, Dong-Kyu Chae
    ICML 2026: Position Papers (Spotlight, Top 5%) (A*) â–Ș â–Ș [PDF]

    As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China's algorithm governance, and the NIST AI Risk Management Framework in the U.S., create a fragmented regulatory landscape. In this position paper, we argue that AI governance must be built not on laws alone, but on ISO-like interoperability protocols that enable standardized, machine-readable risk communication across borders. Drawing on the success of the GDPR, which was operationalized through standards like ISO 27001 and Privacy by Design, we propose the development of standardized AI nutrition labels containing unified metrics for bias, energy usage, and data provenance to facilitate cross-jurisdictional compliance. These manifests would lower barriers for small and medium enterprises (SMEs), reduce redundant regulatory efforts, and build public trust. The paper addresses concerns that standards may stifle innovation by advocating for modular, versioned protocols designed to evolve in tandem with technological change. Overall, we call for a shift from siloed legal compliance toward interoperable technical conformance, enabling a shared global language for responsible AI deployment.
                    @inproceedings{
                    wasi2026position,
                    title={Position: {AI} Governance Needs {ISO}-like Interoperability Protocols, Not Just Laws},
                    author={Azmine Toushik Wasi and Mst Rafia Islam and Mahfuz Ahmed Anik and Taki Hasan Rafi and Md Manjurul Ahsan and Dong-Kyu Chae},
                    booktitle={Forty-third International Conference on Machine Learning Position Paper Track},
                    year={2026},
                    url={https://openreview.net/forum?id=TE3ceHd4YU}
                    }
  3. TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings
    Azmine Toushik Wasi*, Shahriyar Zaman Ridoy*, Koushik Tonmoy, Kinga Tshering, S. M. M. Hasan, Wahid Faisal, T. Mohiuddin, Md Rizwan Parvez
    ICML 2026 (A*)

    Geo-temporal understanding, the ability to identify the location, time, and contextual features of an image from visual cues alone, is a fundamental aspect of human intelligence with wide-ranging applications, from disaster response to navigation and geography education. While recent vision–language models (VLMs) have shown progress in image geo-localization using conspicuous cues like landmarks or road signs, their ability to understand temporal signals and related spatial reasoning cues remains underexplored. To address this gap, we introduce TimeSpot, a comprehensive benchmark for evaluating real-world geo-temporal reasoning in VLMs. TimeSpot contains 1,455 images from 80 countries, where models must infer temporal attributes (season, month, time of day, daylight phase) and geolocation attributes (continent, country, climate zone, environment type, latitude–longitude coordinates) directly from the image. In addition, it includes spatial reasoning tasks that require integrating geographical, spatial, and temporal cues to solve complex understanding problems. Unlike prior benchmarks that emphasize obvious cues or iconic imagery, TimeSpot prioritizes diverse and subtle settings, reflecting the difficulty of reasoning under real-world uncertainty. Our evaluation of state-of-the-art VLMs, including both open- and closed-source models, shows consistently low performance across tasks, underscoring the substantial challenges that remain for robust temporal and geographic reasoning and the need for improved methods to achieve reliable and trustworthy geo-temporal understanding in VLMs.
                    @inproceedings{
                    wasi2026timespot,
                    title={TimeSpot: Benchmarking Geo-Temporal Understanding in Vision{\textendash}Language Models in Real-World Settings},
                    author={Azmine Toushik Wasi and Shahriyar Zaman Ridoy and Koushik Ahamed Tonmoy and Kinga Tshering and S. M. Muhtasimul Hasan and Wahid Faisal and Tasnim Mohiuddin and Md Rizwan Parvez},
                    booktitle={Forty-third International Conference on Machine Learning},
                    year={2026},
                    url={https://openreview.net/forum?id=XQlUqVCHJd}
                    }
  4. SEISMOS: A Statistical Signal Detection Framework for Semantic Chunking
    Ishtiak Mahmud Saad, Mominul Islam, Shahriyar Zaman Ridoy, Md Manjurul Ahsan, Azmine Toushik Wasi$
    NeurIPS 2026: Main Technical Track (A*)

    Chunking is a hidden bottleneck in dense retrieval and retrieval-augmented generation: it determines the units that can be indexed, retrieved, and ultimately used as evidence. Yet most systems still segment documents using fixed token windows or globally thresholded semantic similarity drops, treating chunking as an engineering heuristic rather than a statistical decision. We propose SEISMOS, a statistical signal detection framework for semantic chunking. SEISMOS models consecutive sentence-embedding cosine similarities as a document-level signal and derives boundary decisions from the null hypothesis of no semantic transition. Across four development BEIR corpora, we establish three corpus-invariant properties: bimodal document structure, rapid survival decay of shallow local minima, and positive lag-1 autocorrelation. These properties show that semantic boundaries should be detected by variance-normalized deviations, not by document means or global similarity thresholds. Accounting for autocorrelation leads to a normalized discrete Laplacian detector that identifies significant semantic valleys through a single interpretable decision rule. Evaluated on five BEIR benchmarks, SEISMOS consistently improves over fixed-length, recursive, and production semantic chunking baselines under one fixed operating configuration. Without corpus-specific retuning, the same detector transfers to a held-out TREC-COVID corpus. Our results suggest that the boundaries needed for effective retrieval are already encoded in embedding-similarity signals, and that principled, efficient, LLM-free chunking can be obtained by detecting them statistically.
                    @inproceedings{
                    saad2026seismos,
                    title={{SEISMOS}: A Statistical Signal Detection Framework for Semantic Chunking},
                    author={Ishtiak Mahmud Saad and Mominul Islam and Shahriyar Zaman Ridoy and Md Manjurul Ahsan and Azmine Toushik Wasi},
                    booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
                    year={2026},
                    url={https://openreview.net/forum?id=MiRf0wVbnB}
                    }
  5. From Language Specifications to Executable Turing Machines: Evaluating LLMs as Computational Machine Designers
    Shahriyar Zaman Ridoy, S. M. Muhtasimul Hasan, Azmine Toushik Wasi, Kinga Tshering, Koushik Ahamed Tonmoy, Md Mosaddek Khan
    EMNLP 2026 (Findings, Top 30%) (A*)

    Recognizing a formal language does not imply the ability to construct the computation that decides it. We operationalize this distinction as an executable benchmark for language-model reasoning: given a formal-language specification, an LLM must synthesize a deterministic or nondeterministic Turing machine that decides the language. We introduce TMEval, a benchmark of 253 language specifications across 7 structural categories spanning context-free and non-context-free phenomena. Each language includes 600 oracle-labeled membership tests, and every output is evaluated as an executable artifact rather than a textual response. Our pipeline enforces a strict output contract, validates machine construction, executes each buildable machine, and assigns one of four outcomes: static_fail, build_fail, semantic_mismatch, or exact_pass. Across eight LLMs, 2,024 model-language candidates, and up to 1,214,400 behavioral evaluations, only 11.12% exactly pass all tests. More than half fail before execution, over one-third implement the wrong language, and 75 languages remain unsolved by all systems. Failure depends less on the context-free/non-context-free boundary than on the operational structure required to realize the specification. TMEval reveals a substantial gap between formal language recognition and machine-checked symbolic construction.
                    @inproceedings{
                    anonymous2026from,
                    title={From Language Specifications to Executable Turing Machines: Evaluating {LLM}s as Computational Machine Designers},
                    author={Shahriyar Zaman Ridoy and S. M. Muhtasimul Hasan and Azmine Toushik Wasi and Kinga Tshering and Koushik Ahamed Tonmoy and Md Mosaddek Khan},
                    booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
                    year={2026},
                    url={https://openreview.net/forum?id=tLXgymd1rR}
                    }
  6. Reward Engineering for Software Tasks: A Survey of Reinforcement Learning Approaches
    Md Rayhanul Masud, Azmine Toushik Wasi, Salman Rahman, Md Rizwan Parvez
    EMNLP 2026 (Findings, Top 30%) (A*)

    Reinforcement learning is increasingly used for code-centric software engineering tasks, including code generation, understanding, repair, testing, and optimization, especially with the rise of large language models and autonomous agents. A core challenge in these settings is reward design. Unlike standard RL domains with clear scalar objectives, software tasks involve competing goals such as correctness, security, efficiency, and readability, which are difficult to capture with a single reward. As a result, RL-for-SE systems rely on heterogeneous signals, including compilation results, unit tests, coverage metrics, retrieval scores, and learned preferences. Yet this work remains scattered across tasks and communities. To our knowledge, this survey provides the first systematic review of reward engineering for RL in software tasks. We organize prior work by reward source, granularity, and aggregation. We then distill the findings into practical guidance, including a decision matrix, a design guide, and a reward engineering reporting standard for future RL- for-SE systems.
    @inproceedings{masuds2026reward,
    title={Reward Engineering for Software Tasks: A Survey of Reinforcement Learning Approaches},
    author={Md Rayhanul Masud and Azmine Toushik Wasi and Salman Rahman and Md Rizwan Parvez},
    booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
    year={2026},
    url={https://openreview.net/forum?id=QYatzDTUPu}
    }
  7. PD-scWorld: Pathway-Guided Disentanglement for Single-Cell Perturbation World Models
    Azmine Toushik Wasi
    MLCB 2026 | ACM BCB 2026 (Top 20%, Oral) â–Ș

    Disentangling the latent factors that drive cellular responses remains challenging for generative and predictive models of single-cell data, particularly under perturbations where multiple biological programs co-activate. We propose PD-scWorld, an intervention-aware latent world model in which each latent factor is encouraged to respond selectively to specific perturbations and covariates using weak biological supervision from pathway tags of perturbed genes rather than full factor labels. Given paired pre/post states, the model learns action-conditioned latent transitions while constraining perturbation-induced changes to remain concentrated within latent dimensions associated with the relevant pathway. We introduce a pathway-conditional regularizer that combines group sparsity with a variance-based constraint to localize consistent perturbation effects and suppress entangled drift across unrelated latent dimensions. This produces latents that align with known biological programs while retaining predictive flexibility for unseen perturbations. We evaluate PD-scWorld on Perturb-seq and related CRISPR single-cell screens using gene-to-pathway mappings and cell-cycle annotations, measuring mutual information between latents and covariates, recovery of pathway-specific responses, counterfactual consistency under targeted rollouts, and perturbation-effect prediction. Across datasets, PD-scWorld produces cleaner factorization and more interpretable perturbation mechanisms than ÎČ-VAE and unstructured latent-dynamics baselines, while improving perturbation-effect prediction and robustness to covariate shifts.
    @inproceedings{
    wasi2026pdscworld,
    title={{PD}-scWorld: Pathway-Guided Disentanglement for Single-Cell Perturbation World Models},
    author={Azmine Toushik Wasi},
    booktitle={The 21st Machine Learning in Computational Biology (MLCB)},
    year={2026}
    }
  8. Frugal Medical AI as an Equity Imperative: Rethinking Algorithm Design for Resource-Constrained Healthcare
    Azmine Toushik Wasi, Mohsin Mahmud Topu, Mahfuz Ahmed Anik, Md Manjurul Ahsan
    ACM EAAMO 2026 (Top 22.5%, Oral) â–Ș

    Medical artificial intelligence has achieved high diagnostic accuracy in controlled, resource-rich settings, yet these gains persistently fail to translate into clinical impact across the majority of global healthcare systems. We argue that this failure is not incidental: prevailing medical AI paradigms embed assumptions about high-field imaging, continuous cloud connectivity, and unconstrained computation that are structurally incompatible with the material realities of low- and middle-income country (LMIC) healthcare. This constitutes a form of technical inequity, one enforced by design choices that implicitly treat high-income infrastructure as universal. Positioning the Global South as a stress test for deployment realism, we analyse how four compounding factors, imaging scarcity, absent or informal digitization, environmental constraints, and operator variability, systematically undermine conventional deep learning pipelines. We propose \emph{Frugal Medical AI} as a distinct technical framework that prioritizes edge-native inference, hardware-aware model design, cross-modality generative translation, and robustness to extreme image degradation. We further introduce deployment-aligned evaluation metrics that extend beyond accuracy to incorporate energy consumption, memory access cost, latency, and clinical risk sensitivity. Reframing medical AI innovation around clinical feasibility, environmental sustainability, and global equity, this work argues for a principled set of technical priorities that treat deployment constraints as first-class design objectives rather than downstream concessions.
    @inproceedings{
    wasi2026frugal,
    title={Frugal Medical {AI} as an Equity Imperative: Rethinking Algorithm Design for Resource-Constrained Healthcare},
    author={Azmine Toushik Wasi and Mohsin Mahmud Topu and Mahfuz Ahmed Anik and Md Manjurul Ahsan},
    booktitle={Sixth ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization},
    year={2026},
    url={https://openreview.net/forum?id=6LuDtd3kCC}
    }
  9. Position: Time-Series Foundation Models Require Explicit Domain-Level Benchmarks
    Md Asif Bin Syed, Md Younus Ahamed, Azmine Toushik Wasi$
    ICML 2026: Position Papers (A*)

    Time series foundation models (TSFMs) have demonstrated strong performance on established benchmarks such as GIFT-Eval, Monash, and TSFM-Bench. However, these benchmarks pool datasets from many domains with uneven representation, which can obscure performance within specific application areas such as healthcare, finance, nature, retail, and transport. The necessity for domain-specific evaluation arises from the inherent structural diversity of time series data: clinical records often feature irregular sampling and informative missingness; financial sequences are characterized by high noise and stochastic trajectories; and environmental data, such as energy and weather, are governed by deterministic physical laws and strong seasonal hierarchies. Motivated by this heterogeneity, we argue that TSFMs require explicit domain-specific benchmarks so practitioners can reliably assess a model's utility within their own application area. This is because cross-domain differences in data generation, sampling irregularity, and nonstationarity under concept drift fundamentally shape forecasting difficulty and failure modes. As a result, strong performance on aggregated leaderboards may not translate to reliable deployment within a specific domain. To test this, we evaluated seven TSFMs across 72 datasets from six domains (healthcare, finance, energy, nature, transport, and retail) and found substantial cross-domain variability. These findings confirm that global benchmark scores can be misleading and that domain-aware evaluations are essential for trustworthy TSFM selection.
    @inproceedings{
    syed2026position,
    title={Position: Time-Series Foundation Models Require Explicit Domain-Level Benchmarks},
    author={Md Asif Bin Syed and Md Younus Ahamed and Azmine Toushik Wasi},
    booktitle={Forty-third International Conference on Machine Learning Position Paper Track},
    year={2026},
    url={https://openreview.net/forum?id=W2eEMPjzIQ}
    }
  10. Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice
    Azmine Toushik Wasi, Wahid Faisal, Mst Rafia Islam, Md Rizwan Parvez
    ACL 2026 (A*) Findings â–Ș [arXiv] [ACL Anthology]

    Bangladesh’s low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs. Existing AI legal assistants lack Bengali-language support and jurisdiction-specific adaptation, limiting their effectiveness. To address this, we developed Mina, a multilingual LLM-based legal assistant tailored for the Bangladeshi context. It employs multilingual embeddings and a RAG-based chain-of-tools framework for retrieval, reasoning, translation, and document generation, delivering context-aware legal drafts, citations, and plain-language explanations via an interactive chat interface. Evaluated by law faculty from leading Bangladeshi universities across all stages of the 2022 and 2023 Bangladesh Bar Council examinations, Mina achieved scores of 75–80% in the preliminary MCQs, written, and simulated viva voce components. These results matched or surpassed average human performance, demonstrating strong clarity, contextual understanding, and sound legal reasoning, while operating at approximately 0.1-0.6% of the cost of human lawyers. These results confirm its potential as a low-cost, multilingual AI assistant that automates key legal tasks and scales access to justice, offering a real-world details on building domain-specific, low-resource systems and addressing challenges of multilingual adaptation, efficiency, and sustainable public-service AI deployment.
    @inproceedings{wasi-etal-2026-mina,
        title = "Mina: A Multilingual {LLM}-Powered Legal Assistant Agent for Empowering Access to Justice in {B}angladesh",
        author = "Wasi, Azmine Toushik  and Faisal, Wahid  and Islam, Mst Rafia  and Parvez, Md Rizwan",
        booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
        month = jul,
        year = "2026",
        address = "San Diego, California, United States",
        publisher = "Association for Computational Linguistics",
        url = "https://aclanthology.org/2026.findings-acl.1295/",
        doi = "10.18653/v1/2026.findings-acl.1295",
        pages = "25980--26028",
        ISBN = "979-8-89176-395-1",
    }
  11. CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
    Pedro Ortiz Suarez, Laurie Burchell, ..., Azmine Toushik Wasi, ..., Kenton Murray, Sarah K. K. Luger
    ACL 2026 (A*) Main â–Ș [arXiv] [ACL Anthology]

    Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID’s value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.
    @inproceedings{suarez-etal-2026-commonlid,
        title = "{C}ommon{LID}: Re-evaluating State-of-the-Art Language Identification Performance on Web Data",
        author = "Suarez, Pedro Ortiz  and Burchell, Laurie  and Arnett, Catherine  and Mosquera, Rafael  and Monsalve, Sara Hincapi{\'e}  and Vaughan, Thom  and Stewart, Damian  and Ostendorff, Malte  and Abdulmumin, Idris  and Marivate, Vukosi  and Muhammad, Shamsuddeen Hassan  and Tonja, Atnafu Lambebo  and Al-Khalifa, Hend  and Hammouda, Nadia Ghezaiel  and Otiende, Verrah Akinyi  and Wong, Tack Hwa  and Saydaliev, Jakhongir  and Nobakhtian, Melika  and Habibi, Muhammad Ravi Shulthan  and Kranti, Chalamalasetti  and Muchemi, Carol  and Nguyen, Khang  and Adam, Faisal Muhammad  and Salim, Luis Frentzen  and Alqifari, Reem  and Amol, Cynthia Jayne  and Imperial, Joseph Marvin  and Kesen, Ilker  and Mustafid, Ahmad  and Stepachev, Pavel  and Choshen, Leshem  and Anugraha, David  and Nayel, Hamada  and Yimam, Seid Muhie  and Alexandra Putra, Vallerie  and Nguyen, My Chiffon  and Wasi, Azmine Toushik  and Vadithya, Gouthami  and van der Goot, Rob  and C{'}horr, Lanwenn ar  and Dua, Karan  and Yates, Andrew  and Bangera, Mithil  and Bangera, Yeshil  and Patel, Hitesh Laxmichand  and Okabe, Shu  and Ilasariya, Fenal Ashokbhai  and Gaynullin, Dmitry  and Winata, Genta Indra  and Li, Yiyuan  and Mart{\'i}nez, Juan Pablo  and Agarwal, Amit  and Hanif, Ikhlasul Akmal  and Ahmad, Raia Abu  and Adenuga, Esther  and Tjiaranata, Filbert Aurelian  and Buaphet, Weerayut  and Anugraha, Michael  and Vajjala, Sowmya  and Rice, Benjamin L  and Amirudin, Azril Hafizi  and Alabi, Jesujoba Oluwadara  and Panda, Srikant  and Toughrai, Yassine  and Kyomuhendo, Bruhan  and Ruffinelli, Daniel  and Akshata  and Goul{\~a}o, Manuel  and Zhou, Ej  and Ramirez, Ingrid Gabriela Franco  and Aggazzotti, Cristina  and Dobler, Konstantin  and Kevin, Jun  and Pag{\`e}s, Quentin  and Andrews, Nicholas  and Ibrahim, Nuhu  and Ruckdeschel, Mattes  and Keleg, Amr  and Zhang, Mike  and Muziri, Casper Rufaro  and Samuel, Saron  and Takeshita, Sotaro  and Kerdthaisong, Kun  and Foppiano, Luca  and Dent, Rasul  and Green, Tommaso  and Wali, Ahmad Mustapha  and Makaaka, Kamohelo  and Feliren, Vicky  and Idris, Inshirah  and Celikkanat, Hande  and Abubakar, Abdulhamid  and Maillard, Jean  and Sagot, Beno{\^i}t  and Cl{\'e}rice, Thibault  and Murray, Kenton  and Luger, Sarah K. K.",
        booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 1: Long Papers)",
        month = jul,
        year = "2026",
        address = "San Diego, California, United States",
        publisher = "Association for Computational Linguistics",
        url = "https://aclanthology.org/2026.acl-long.1527/",
        doi = "10.18653/v1/2026.acl-long.1527",
        pages = "33063--33080",
        ISBN = "979-8-89176-390-6",
    }
  12. Human-AI Interaction in Low-Tech, Low-Resource Healthcare Systems
    Azmine Toushik Wasi, Mahdiya Rahman Sukanya
    ACM Interactive Health 2026

    Low-resource healthcare systems in South Asia and Africa face severe constraints in infrastructure, connectivity, data availability, and digital literacy that shape how artificial intelligence can be deployed. This paper examines human–AI interaction in these settings, focusing on patients, community health workers, and clinicians rather than model performance alone. Drawing on empirical studies and case examples from Bangladesh, India, and Nigeria, it shows that AI systems designed for high-income contexts often fail when transferred without adaptation. Key challenges include language diversity, absence of electronic health records, limited AI literacy, trust deficits, and unresolved ethical and liability concerns. The analysis demonstrates that localized, multilingual, human-in-the-loop AI can meaningfully augment care when integrated into existing workflows and mediated by trusted health workers. Overall, the paper synthesizes design principles for human-centered AI that emphasize localization, explainability, training, and accountability, arguing that effectiveness depends more on interaction design than technical sophistication.
    @inproceedings{10.1145/3786579.3804944,
    author = {Wasi, Azmine Toushik and Sukanya, Mahdiya Rahman},
    title = {Human--AI Interaction in Low-Tech, Low-Resource Healthcare Systems},
    year = {2026},
    isbn = {9798400724220},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3786579.3804944},
    doi = {10.1145/3786579.3804944},
    abstract = {Low-resource healthcare systems in South Asia and Africa face severe constraints in infrastructure, connectivity, data availability, and digital literacy that shape how artificial intelligence can be deployed. This paper examines human–AI interaction in these settings, focusing on patients, community health workers, and clinicians rather than model performance alone. Drawing on empirical studies and case examples from Bangladesh, India, and Nigeria, it shows that AI systems designed for high-income contexts often fail when transferred without adaptation. Key challenges include language diversity, absence of electronic health records, limited AI literacy, trust deficits, and unresolved ethical and liability concerns. The analysis demonstrates that localized, multilingual, human-in-the-loop AI can meaningfully augment care when integrated into existing workflows and mediated by trusted health workers. Overall, the paper synthesizes design principles for human-centered AI that emphasize localization, explainability, training, and accountability, arguing that effectiveness depends more on interaction design than technical sophistication.},
    booktitle = {Proceedings of the 2026 ACM Interactive Health Conference},
    articleno = {38},
    numpages = {7},
    keywords = {Human–AI Interaction, Low-Resource Environments, Low-Technology Environments, Global South, AI In Healthcare, Community Health Workers, Multilingual Interfaces, Health AI Design},
    location = {
    },
    series = {IH '26}
    }
  13. Clinician Perspectives on Generative AI Integration in Bangladesh's Healthcare Workflows: A Brief Study
    Azmine Toushik Wasi, Mahir Absar Khan, Rahatun Nesa Priti, Abdur Rahman, Mst Rafia Islam, Miskat Tahmin Ali
    ACM Interactive Health 2026

    Generative artificial intelligence is increasingly proposed as a means to support clinical decision-making, documentation, and care coordination, yet its integration into everyday healthcare workflows remains challenging in low-resource settings. This paper examines clinician perspectives on integrating generative AI into healthcare workflows in Bangladesh, where infrastructure limitations, data fragmentation, and workforce shortages shape technology adoption. We conducted semi-structured interviews with eleven healthcare professionals, including physicians, academic medical experts, nurses, and support staff, to understand clinical practices, workflow bottlenecks, and expectations around generative AI. Findings reveal persistent challenges such as incomplete patient histories, high patient volumes, limited access to diagnostics, fragmented data systems, and administrative burden contributing to clinician burnout. Participants identified potential roles for generative AI in record summarization, triage support, documentation, and treatment planning, while expressing concerns related to data privacy, reliability, cost, and contextual fit. We discuss implications for human-centered design that complement clinical expertise in overburdened healthcare environments.
    @inproceedings{10.1145/3786579.3804935,
    author = {Wasi, Azmine Toushik and Khan, Mahir Absar and Priti, Rahatun Nesa and Rahman, Abdur and Islam, Mst Rafia and Ali, Miskat Tahmin},
    title = {Clinician Perspectives on Generative AI Integration in Bangladesh’s Healthcare Workflows: A Brief Study},
    year = {2026},
    isbn = {9798400724220},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3786579.3804935},
    doi = {10.1145/3786579.3804935},
    abstract = {Generative artificial intelligence is increasingly proposed as a means to support clinical decision-making, documentation, and care coordination, yet its integration into everyday healthcare workflows remains challenging in low-resource settings. This paper examines clinician perspectives on integrating generative AI into healthcare workflows in Bangladesh, where infrastructure limitations, data fragmentation, and workforce shortages shape technology adoption. We conducted semi-structured interviews with eleven healthcare professionals, including physicians, academic medical experts, nurses, and support staff, to understand clinical practices, workflow bottlenecks, and expectations around generative AI. Findings reveal persistent challenges such as incomplete patient histories, high patient volumes, limited access to diagnostics, fragmented data systems, and administrative burden contributing to clinician burnout. Participants identified potential roles for generative AI in record summarization, triage support, documentation, and treatment planning, while expressing concerns related to data privacy, reliability, cost, and contextual fit. We discuss implications for human-centered design that complement clinical expertise in overburdened healthcare environments.},
    booktitle = {Proceedings of the 2026 ACM Interactive Health Conference},
    articleno = {8},
    numpages = {6},
    keywords = {Clinical Decision Support, Low-Resource Settings, AI In Healthcare, Digital Health Systems, Low-Technology Environments},
    location = {
    },
    series = {IH '26}
    }
  14. Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
    Israfel Salazar, Manuel FernĂĄndez Burda, Shayekh Bin Islam, ..., Azmine Toushik Wasi, ..., Sara Hooker, and Marzieh Fadaee
    ICLR 2026 (A*) â–Ș [PDF]

    The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While multilingual benchmarks have expanded, both in size and language, many rely on translations of English datasets, failing to capture cultural nuances. In this work, we propose Kaleidoscope, as the most comprehensive exam benchmark to date for the multilingual evaluation of vision-language models. Kaleidoscope is a large-scale, in-language multimodal benchmark designed to evaluate VLMs across diverse languages and visual inputs. Kaleidoscope covers 18 languages and 14 different subjects, amounting to a total of 20,911 multiple-choice questions. Built through an open science collaboration with a diverse group of researchers worldwide, Kaleidoscope ensures linguistic and cultural authenticity. We evaluate top-performing multilingual vision-language models and find that they perform poorly on low-resource languages and in complex multimodal scenarios. Our results highlight the need for progress on culturally inclusive multimodal evaluation frameworks.
    @inproceedings{
    salazar2026kaleidoscope,
    title={Kaleidoscope: In-language Exams for Massively  Multilingual Vision Evaluation},
    author={Israfel Salazar and Manuel Fern{\'a}ndez Burda and Shayekh Bin Islam and Arshia Soltani Moakhar and Shivalika Singh and Fabian Farestam and Angelika Romanou and Danylo Boiko and Dipika Khullar and Mike Zhang and Dominik Krzemi{\'n}ski and Jekaterina Novikova and Lu{\'\i}sa Shimabucoro and Joseph Marvin Imperial and Rishabh Maheshwary and Sharad Duwal and Alfonso Amayuelas and Swati Rajwal and Jebish Purbey and Ahmed Ruby and Nicholas Popovi{\v{c}} and Marek Suppa and Azmine Toushik Wasi and Ram Mohan Rao Kadiyala and Olga Tsymboi and Maksim Kostritsya and Bardia soltani moakhar and Gabriel da Costa Merlin and Ot{\'a}vio Ferracioli Coletti and Maral Jabbarishiviari and MOHAMMADAMIN FARAHANIFARD and Silvia Andrea Fernandez and Mar{\'\i}a Grandury and Dmitry Abulkhanov and Drishti Sharma and Andre Guarnier De Mitri and Leticia Bossatto Marchezi and Setayesh Heydari and Johan Obando-Ceron and Nazar Kohut and Beyza Ermis and Desmond Elliott and Enzo Ferrante and Sara Hooker and Marzieh Fadaee},
    booktitle={The Fourteenth International Conference on Learning Representations},
    year={2026},
    url={https://openreview.net/forum?id=zCYXhSy9UH}
    }
  15. Can an AI Teach Children to Recognize Medical Emergencies? A Game-Based Human-AI Learning Approach
    Azmine Toushik Wasi, Mst Rafia Islam, Wahid Faisal, Abdur Rahman
    AIED 2026 (A) â–Ș Late Breaking Results (LBR: Posters)

    Children are often the first witnesses to medical emergencies at home but lack the knowledge or confidence to respond effectively. We present an AI-powered game designed to teach children (ages 7–12) to recognize signs of medical emergencies (e.g., stroke, seizure, unconsciousness) and practice appropriate actions in a safe virtual environment. The system features a scenario generation engine with progressive difficulty, speech-enabled 911 call practice, and real-time feedback on recognition accuracy. In a pilot study, children played through emergency scenarios, and we evaluated their ability to identify emergencies before and after gameplay. Preliminary results indicate improved recognition and preparedness, while parental interviews suggest the game is both empowering and educational without causing undue fear. This work highlights the potential of AI as an interactive learning teammate in high-stakes safety education.
    @inbook{Wasi2026,
      title = {Can an AI Teach Children to Recognize Medical Emergencies? A Game-Based Human-AI Learning Approach},
      ISBN = {9783032297884},
      ISSN = {1865-0937},
      url = {http://dx.doi.org/10.1007/978-3-032-29788-4_5},
      DOI = {10.1007/978-3-032-29788-4_5},
      booktitle = {Artificial Intelligence in Education. Late Breaking Results,  WideAIED,  Practitioners,  Industry and Policies,  Blue Sky,  Doctoral Consortium,  FoL Workshops and Tutorials,  FoL Invited Papers},
      publisher = {Springer Nature Switzerland},
      author = {Wasi,  Azmine Toushik and Islam,  Mst Rafia and Faisal,  Wahid and Rahman,  Abdur},
      year = {2026},
      pages = {31–37}
    }
  16. EmphAI: An VLM-Based System for Feedback on Non-Verbal Communication in Clinical Training
    Azmine Toushik Wasi, Wahid Faisal, Mst Rafia Islam
    AIED 2026 (A) â–Ș Late Breaking Results (LBR: Posters)

    Non-verbal communication (NVC)—including eye contact, posture, facial expression, and use of silence—is critical yet difficult to teach in clinical education, with traditional feedback from standardized patients and instructors limited in scalability. We introduce EmphAI, a vision-language model (VLM)-based system that analyzes sampled video frames of student–patient interactions to generate structured, timestamped feedback on NVC using text-only few-shot prompting, avoiding complex multimodal pipelines. We evaluate EmphAI through a quantitative user study with 53 trainees assessing usability and perceived collaboration, and a pre/post study with 24 participants measuring changes in self-reported NVC confidence. Results indicate high usability and usefulness, with participants reporting clear, actionable feedback that supports reflective practice, and a statistically significant increase in confidence across multiple NVC dimensions. It suggest that VLM-based systems can serve as scalable, supportive feedback tools, providing early evidence for multimodal AI as a collaborative teammate in medical education.
    @inbook{Wasi2026,
      title = {EmphAI: An VLM-Based System for Feedback on Non-Verbal Communication in Clinical Training},
      ISBN = {9783032297884},
      ISSN = {1865-0937},
      url = {http://dx.doi.org/10.1007/978-3-032-29788-4_4},
      DOI = {10.1007/978-3-032-29788-4_4},
      booktitle = {Artificial Intelligence in Education. Late Breaking Results,  WideAIED,  Practitioners,  Industry and Policies,  Blue Sky,  Doctoral Consortium,  FoL Workshops and Tutorials,  FoL Invited Papers},
      publisher = {Springer Nature Switzerland},
      author = {Wasi,  Azmine Toushik and Faisal,  Wahid and Islam,  Mst Rafia},
      year = {2026},
      pages = {24–30}
    }
  17. Spatial Reasoning in Multimodal Foundation Models: A Survey of Spatial Understanding in LLMs and VLMs
    Azmine Toushik Wasi, Md. Masudur Rahman, Abdur Rahman, Mohsin Mahmud Topu, Mohammed Eunus Ali, Riashat Islam, Md Rizwan Parvez
    AACL-IJCNLP 2026: Main (B)

    Spatial reasoning, the ability to represent and reason about spatial relations, geometry, topology, and embodied environments, is increasingly vital for evaluating and deploying large language models (LLMs) and vision–language models (VLMs). While multimodal foundation models show strong performance on spatial benchmarks, navigation, and embodied planning, much of this stems from shallow pattern matching rather than grounded spatial understanding. This review offers a comprehensive, framework-driven analysis of spatial reasoning in LLMs and VLMs, covering tasks, representations, datasets, benchmarks, architectures, training, and failure modes. We also link spatial reasoning to embodied AI, emphasizing challenges in generalization, compositionality, and grounding. Our survey unifies diverse fields, NLP, vision, robotics, and cognitive science, to guide the development of truly spatially grounded foundation models.
  18. Geo-spatial and Geo-temporal Reasoning in Vision-Language and Large Language Models: A Review
    Shahriyar Zaman Ridoy, Azmine Toushik Wasi, Koushik Ahamed Tonmoy, Md Rizwan Parvez
    AACL-IJCNLP 2026: Main (B)

    Geo-spatial and geo-temporal reasoning require models to determine not only what information is present but also where and when it applies, which is critical for applications such as crisis response, environmental monitoring, and urban planning. This review surveys recent advances in vision–language models (VLMs) and large language models (LLMs) that address geographically and temporally grounded reasoning across tasks including image geolocation, spatiotemporal question answering, language-guided navigation, and Earth observation. We organize the literature through a structured analysis covering task families, input modalities, model architectures, and reasoning paradigms, highlighting how perception, retrieval, tool use, and intermediate reasoning steps contribute to system performance. The survey further synthesizes existing datasets and benchmarks, identifying common evaluation pitfalls such as spatial granularity mismatch, temporal drift, geographic bias, and potential data leakage. Finally, we outline key research challenges and opportunities for developing robust, verifiable, and time-aware geo-spatial AI systems while addressing broader concerns related to fairness, bias, and location privacy.
  19. Large Language and Vision-Language Models as World Models: A Survey
    Azmine Toushik Wasi, Sadia Tasnim Meem, Riashat Islam, Md Rizwan Parvez
    AACL-IJCNLP 2026: Findings (B)

    Recent advances have reframed Artificial Intelligence (AI) around world models, internal simulators that predict environmental dynamics and support general intelligence. This review surveys Large Language Models (LLMs) and Vision–Language Models (VLMs) not merely as text or vision processors, but as predictive systems capable of modeling state, causality, and dynamics. We introduce a dual taxonomy to organize post-2021 research: a theoretical taxonomy distinguishing pixel-space generative simulators, latent predictive architectures such as JEPA, and neuro-symbolic hybrids; and an application taxonomy covering autonomous driving, embodied robotics, interactive gaming, and social simulation. Synthesizing recent works, we conclude that although autoregressive transformers exhibit emergent topological and causal reasoning, robust physical world modeling requires moving beyond next-token prediction toward hybrid architectures that integrate generative diffusion models with symbolic grounding.
  20. Is There a Dr. House in the Machine? A Survey of Clinical and Diagnostic Reasoning in LLM-based Medical Agents
    Azmine Toushik Wasi, Abdur Rahman, Riashat Islam, Md Rizwan Parvez
    AACL-IJCNLP 2026: Findings (B)

    Large Language Models (LLMs) are increasingly embedded into agentic systems capable of tool use, multi-step reasoning, and interaction with clinical knowledge sources. These systems promise to move beyond pattern matching toward diagnostic reasoning resembling that of expert clinicians. This survey provides a comprehensive review of LLM-based medical agents, focusing on diagnostic reasoning, tool augmentation, and clinical decision-making. We introduce a unified framework covering reasoning paradigms, agent architectures, tool ecosystems, evaluation methodologies, safety considerations, and deployment challenges. By synthesizing evidence across machine learning, biomedical informatics, and clinical AI literature, we assess whether contemporary systems meaningfully approximate expert diagnostic reasoning or merely simulate it. We conclude by outlining open research challenges and proposing directions toward trustworthy, clinically grounded medical agents.
  21. Position: Real-World Clinical NLP Robustness Requires Messy, Multimodal, Longitudinal, Privacy-Preserving Corpora
    Azmine Toushik Wasi, Shahriyar Zaman Ridoy
    AACL-IJCNLP 2026: Findings (B)

    Real-world healthcare data is inherently complex: noisy, incomplete, heterogeneous, and temporally irregular, making clinically robust AI development difficult. Traditional models, trained on idealized datasets, fail to capture these operational realities, limiting real-world performance. We propose the "Messy Clinic" paradigm: a dataset and blueprint for building AI systems grounded in authentic clinical complexity. This framework enables seamless multimodal data integration across EHRs, imaging, genomics, and wearables, using privacy-preserving methods such as federated learning, synthetic data, and differential privacy. It provides (1) structured methodologies for longitudinal multimodal management, (2) AI techniques resilient to noise and missing data, and (3) governance mechanisms for consent, accountability, and bias mitigation. By embracing the “Messy Clinic,” healthcare AI can shift from artificial idealism to real-world robustness, accelerating research, enabling personalized care, and aligning AI with the true complexity of clinical practice.
  22. A Review of Human-Centric Evaluation of Cultural Biases in Indic Languages in LLMs: Rethinking Research Directions
    Azmine Toushik Wasi, Omid Reza Heidari*, MD Mohaymen Ul Anam*, Abdur Rahman*, Amit Agarwal, Hitesh Laxmichand Patel, Bhargava Kumar, TH Rafi, Dong-Kyu Chae
    PAKDD 2026: DSFA (B) â–Ș [PDF]

    Indic languages, spoken by more than 1.3 billion people in South Asia, are often underrepresented in AI systems due to their diversity, limited resources, and evaluation challenges. This region, known for its many coexisting developed and endangered cultures, faces significant cultural bias risks as these cultures are not equally represented. These biases can significantly impact the design, development, and deployment of AI in real-life scenarios. Our survey explores cultural bias in Indic languages from a human-centered perspective, combining computational social science and NLP. We interpret different types of bias, critically review different evaluation and mitigation works, and identify research gaps and priorities to focus on to improve AI systems. Based on our review findings, we argue that we need to rethink and reprioritize our research efforts to support other Indic languages except Hindi, and focus on different evaluation approaches for cultural biases. We discuss human-centered perspectives to enhance LLMs, focusing on how they can operate more effectively and equitably within South Asia’s diverse linguistic and rich cultural landscape, thereby creating fairer and more inclusive AI systems for real-world applications.
    @inproceedings{10.1007/978-981-92-1947-6_48,
    author = {Wasi, Azmine Toushik and Heidari, Omid Reza and Anam, MD Mohaymen Ul and Rahman, Abdur and Agarwal, Amit and Patel, Hitesh Laxmichand and Kumar, Bhargava and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {A Review of Human-Centric Evaluation of Cultural Biases in Indic Languages in LLMs: Rethinking Research Directions},
    year = {2026},
    isbn = {978-981-92-1946-9},
    publisher = {Springer-Verlag},
    address = {Berlin, Heidelberg},
    url = {https://doi.org/10.1007/978-981-92-1947-6_48},
    doi = {10.1007/978-981-92-1947-6_48},
    abstract = {Indic languages, spoken by more than 1.3 billion people in South Asia, are often underrepresented in AI systems due to their diversity, limited resources, and evaluation challenges. This region, known for its many coexisting developed and endangered cultures, faces significant cultural bias risks as these cultures are not equally represented. These biases can significantly impact the design, development, and deployment of AI in real-life scenarios. Our survey explores cultural bias in Indic languages from a human-centered perspective, combining computational social science and NLP. We interpret different types of bias, critically review different evaluation and mitigation works, and identify research gaps and priorities to focus on to improve AI systems. Based on our review findings, we argue that we need to rethink and reprioritize our research efforts to support other Indic languages except Hindi, and focus on different evaluation approaches for cultural biases. We discuss human-centered perspectives to enhance LLMs, focusing on how they can operate more effectively and equitably within South Asia’s diverse linguistic and rich cultural landscape, thereby creating fairer and more inclusive AI systems for real-world applications.},
    booktitle = {Data Science: Foundations and Applications: 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2026, Hong Kong, China, June 9–12, 2026, Proceedings, Part II},
    pages = {618–636},
    numpages = {19},
    keywords = {Indic Languages, Cultural Bias, Human-Centered NLP, Large Language Models, Model Evaluation and Auditing},
    location = {Hong Kong, China}
    }
  23. Generative Artificial Intelligence for Project Management: A Review of Architectures, Applications, and Governance
    Mahfuz Ahmed Anik*, Azmine Toushik Wasi*, Mohsin Mahmud Topu, TH Rafi, Dong-Kyu Chae
    PAKDD 2026: DSFA (B) â–Ș [PDF]

    Generative artificial intelligence is reshaping contemporary project management (PM) through advances in Large Language Models (LLMs), multimodal foundation models, and autonomous AI agents. Persistent cost and schedule overruns across IT, construction, and infrastructure projects highlight the limitations of traditional PM frameworks and motivate the adoption of adaptive, data-driven decision systems. Between 2020 and 2026, generative AI expanded beyond predictive analytics to support document understanding, long-horizon planning, automated reasoning, and tool-based workflow execution. This review synthesizes evidence from empirical research, industrial applications, and reproducible preprints to evaluate how generative AI enhances planning, scheduling, risk and cost estimation, resource allocation, progress monitoring, document analysis, and stakeholder communication. We propose a unified technical taxonomy covering LLMs, multimodal models, retrieval-augmented systems, and autonomous agents, and evaluate their empirical performance relative to classical machine learning approaches. The review further identifies socio-technical challenges related to hallucination, automation bias, data governance, and human–AI coordination. Collectively, these insights articulate the emerging landscape of generative AI–enabled project management and outline pathways for safe, scalable, and high-impact adoption across complex project ecosystems.
    @inproceedings{10.1007/978-981-92-1947-6_47,
    author = {Anik, Mahfuz Ahmed and Wasi, Azmine Toushik and Topu, Mohsin Mahmud and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {Generative Artificial Intelligence for Project Management: A Review of Architectures, Applications, and Governance},
    year = {2026},
    isbn = {978-981-92-1946-9},
    publisher = {Springer-Verlag},
    address = {Berlin, Heidelberg},
    url = {https://doi.org/10.1007/978-981-92-1947-6_47},
    doi = {10.1007/978-981-92-1947-6_47},
    abstract = {Generative artificial intelligence is reshaping contemporary project management (PM) through advances in Large Language Models (LLMs), multimodal foundation models, and autonomous AI agents. Persistent cost and schedule overruns across IT, construction, and infrastructure projects highlight the limitations of traditional PM frameworks and motivate the adoption of adaptive, data-driven decision systems. Between 2020 and 2026, generative AI expanded beyond predictive analytics to support document understanding, long-horizon planning, automated reasoning, and tool-based workflow execution. This review synthesizes evidence from empirical research, industrial applications, and reproducible preprints to evaluate how generative AI enhances planning, scheduling, risk and cost estimation, resource allocation, progress monitoring, document analysis, and stakeholder communication. We propose a unified technical taxonomy covering LLMs, multimodal models, retrieval-augmented systems, and autonomous agents, and evaluate their empirical performance relative to classical machine learning approaches. The review further identifies socio-technical challenges related to hallucination, automation bias, data governance, and human–AI coordination. Collectively, these insights articulate the emerging landscape of generative AI–enabled project management and outline pathways for safe, scalable, and high-impact adoption across complex project ecosystems.},
    booktitle = {Data Science: Foundations and Applications: 30th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2026, Hong Kong, China, June 9–12, 2026, Proceedings, Part II},
    pages = {599–617},
    numpages = {19},
    keywords = {Generative AI, Large Language Models, AI Agents, Project Management Automation, Multimodal Foundation Models},
    location = {Hong Kong, China}
    }
  24. BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
    Shahriyar Zaman Ridoy*, Azmine Toushik Wasi*, Koushik Ahamed Tonmoy, TH Rafi, Dong-Kyu Chae
    ACM FAccT 2026 â–Ș [ACM DL]

    As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, spoken by over 285 million people worldwide and among the most widely spoken languages globally, remains underexplored. Existing ethics benchmarks are predominantly English-centric and shaped by Western moral frameworks, overlooking cultural nuances vital for real-world deployment. To address this gap, we introduce BengaliMoralBench, a large-scale ethics benchmark designed for Bengali language and sociocultural contexts. Our benchmark spans five moral domains: (1) Daily Activities, (2) Habits, (3) Parenting, (4) Family Relationships, and (5) Religious Activities, each subdivided into ten culturally grounded categories, totaling 50 subtopics. Each scenario is annotated through native-speaker consensus under three ethical lenses: virtue ethics, commonsense ethics, and justice ethics. We conduct a systematic zero-shot evaluation under a unified prompting protocol across both open-weight and closed-source models, including recent Llama and Gemma variants, Qwen and DeepSeek models, frontier models (GPT-4o-mini and Gemini 1.5 Pro), and a large multilingual baseline (Qwen3-Next-80B). Results show substantial variation in performance across lenses and domains, and our qualitative analysis reveals persistent weaknesses in cultural grounding, commonsense reasoning, and moral fairness. These findings expose critical limitations of current LLMs in non-Western settings and underscore the need for culturally grounded evaluation. BengaliMoralBench provides a foundation for responsible localization and benchmarking to support the deployment of language technologies in culturally diverse, low-resource markets such as Bangladesh. BengaliMoralBench is available at: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.
    @inproceedings{10.1145/3805689.3812230,
    author = {Ridoy, Shahriyar Zaman and Wasi, Azmine Toushik and Tonmoy, Koushik Ahamed and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture},
    year = {2026},
    isbn = {9798400725968},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3805689.3812230},
    doi = {10.1145/3805689.3812230},
    abstract = {As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, spoken by over 285 million people worldwide and among the most widely spoken languages globally, remains underexplored. Existing ethics benchmarks are predominantly English-centric and shaped by Western moral frameworks, overlooking cultural nuances vital for real-world deployment. To address this gap, we introduce BengaliMoralBench, a large-scale ethics benchmark designed for Bengali language and sociocultural contexts. Our benchmark spans five moral domains: (1) Daily Activities, (2) Habits, (3) Parenting, (4) Family Relationships, and (5) Religious Activities, each subdivided into ten culturally grounded categories, totaling 50 subtopics. Each scenario is annotated through native-speaker consensus under three ethical lenses: virtue ethics, commonsense ethics, and justice ethics. We conduct a systematic zero-shot evaluation under a unified prompting protocol across both open-weight and closed-source models, including recent Llama and Gemma variants, Qwen and DeepSeek models, frontier models (GPT-4o-mini and Gemini 1.5 Pro), and a large multilingual baseline (Qwen3-Next-80B). Results show substantial variation in performance across lenses and domains, and our qualitative analysis reveals persistent weaknesses in cultural grounding, commonsense reasoning, and moral fairness. These findings expose critical limitations of current LLMs in non-Western settings and underscore the need for culturally grounded evaluation. BengaliMoralBench provides a foundation for responsible localization and benchmarking to support the deployment of language technologies in culturally diverse, low-resource markets such as Bangladesh. BengaliMoralBench is available at: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.},
    booktitle = {Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency},
    pages = {8051–8088},
    numpages = {38},
    location = {
    },
    series = {FAccT '26}
    }
  25. Multimodal Vision-Language Models for Automated and Explainable Postoperative Complication Risk Stratification
    Azmine Toushik Wasi, Mahfuz Ahmed Anik, Md Shafikul Islam, Md Manjurul Ahsan
    IISE 2026 (Health Systems Track > Best Track Paper : Second Place) â–Ș [PDF]

    Postoperative patient deterioration remains a major challenge in intensive care, often going undetected because of alert fatigue and the cognitive burden of integrating diverse clinical data streams. Vital signs and narrative clinical notes are typically reviewed separately, requiring clinicians to manually connect numeric trends with textual observations under time pressure. Existing predictive models generally rely either on structured physiological data alone or on clinical text without grounding in concurrent physiological dynamics, while many lack transparent, evidence-linked outputs suitable for clinical deployment. To address this gap, we propose GROVE (GROunded Vision–languagE risk model for ICU), a multimodal vision-language framework for early postoperative risk detection. GROVE provides explainable and clinically actionable assessments by jointly modeling quantitative physiological signals and qualitative narrative information during the first 48 h following surgery using MIMIC-IV ICU data. The framework integrates: (1) structured physiological measurements, including heart rate, mean arterial pressure, respiratory rate, oxygen saturation, and temperature, represented as temporal plot images; and (2) unstructured nursing and provider notes. Unlike conventional classifiers, GROVE generates structured natural language assessments linking physiological trends with supporting clinical observations. Evaluation on the MIMIC-IV postoperative ICU cohort shows that GROVE achieves an AUROC of 0.92 for early detection of complications such as sepsis and hemorrhage, while 94% of generated outputs are rated evidence-aligned under clinician adjudication. By bridging quantitative and qualitative patient signals within an interpretable multimodal framework, GROVE reduces clinician cognitive burden and supports timely intervention, providing a scalable foundation for predictive and actionable clinical decision support.
    @article{Wasi2026,
      title = {Multimodal vision-language models for automated and explainable postoperative complication risk stratification},
      ISSN = {2472-5587},
      url = {http://dx.doi.org/10.1080/24725579.2026.2703103},
      DOI = {10.1080/24725579.2026.2703103},
      journal = {IISE Transactions on Healthcare Systems Engineering},
      publisher = {Informa UK Limited},
      author = {Wasi,  Azmine Toushik and Anik,  Mahfuz Ahmed and Islam,  MD Shafikul and Ahsan,  Md Manjurul},
      year = {2026},
      month = July,
      pages = {1–16}
    }
  26. VLM-Based Zero-Shot Anomaly Detection and Root-Cause Interpretation in Complex Industrial Assemblies
    Mahfuz Ahmed Anik, Azmine Toushik Wasi, Md Shafikul Islam, Md Manjurul Ahsan
    IISE 2026

    Modern industrial quality inspection systems rely heavily on deep learning models that require extensive labeled defect datasets. However, for complex, low-volume products such as aerospace assemblies or medical devices, defective samples are scarce, making conventional supervised training infeasible. This scarcity forms a critical research gap in developing scalable and interpretable defect detection systems. Moreover, existing approaches typically output only bounding boxes without interpretable explanations, leaving engineers to manually perform rootcause diagnosis a major bottleneck in corrective action workflows. To address these challenges, this work proposes a Vision-Language Model (VLM)-driven, zero-shot anomaly detection framework that reframes defect detection as a Visual Question Answering (VQA) problem. Using pre-trained multimodal models such as GPT-4V or LLaVA, the system analyzes product images through a sequence of structured prompts informed by design specifications and component hierarchies. Each prompt targets global, component-level, and relational checks, producing textual outputs describing anomalies, their nature, and spatial context. A lightweight parser converts these natural language responses into standardized non-conformance reports. Evaluated on the MVTec AD dataset, the proposed system achieves F1score improvements of 18–24% over baseline zero-shot vision models without any task-specific fine-tuning. The generated textual descriptions align with true defect regions in 92% of cases, demonstrating high interpretability and diagnostic value. This approach bridges the gap between perception and reasoning in visual inspection, enabling interpretable, data-efficient, and generalizable defect analysis. Beyond manufacturing, the framework paves the way for autonomous inspection systems capable of human-level reasoning across diverse industrial domains.
  27. Exploring Bengali Creative Storytelling Capabilities of Large Language Models Across Cultural Variations
    Azmine Toushik Wasi, Raima Islam*, Mst Rafia Islam*, Farig Sadeque, TH Rafi, Dong-Kyu Chae
    CSCW 2025 Posters (A) â–Ș [PDF]

    Large Language Models (LLMs) excel in fluency but often struggle with originality, suspense, and emotional depth in storytelling. This study evaluates their creative storytelling capabilities in Bengali, a language with significant dialectal diversity. Using three narrative prompts across single-dialect and cross-dialect settings with initial results and story continuation, we analyze AI-generated content for coherence, creativity, and cultural relevance. Native Bengali speakers provide qualitative feedback, highlighting key challenges such as dialectal fidelity and narrative richness. Our findings emphasize the need for culturally adaptive NLP models to enhance storytelling in low-resource languages.
    @inproceedings{10.1145/3715070.3749228,
    author = {Wasi, Azmine Toushik and Islam, Raima and Islam, Mst Rafia and Sadeque, Farig Y. and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {Exploring Bengali Creative Storytelling Capabilities of Large Language Models Across Cultural Variations},
    year = {2025},
    isbn = {9798400714801},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3715070.3749228},
    doi = {10.1145/3715070.3749228},
    abstract = {Large Language Models (LLMs) excel in fluency but often struggle with originality, suspense, and emotional depth in storytelling. This study evaluates their creative storytelling capabilities in Bengali, a language with significant dialectal diversity. Using three narrative prompts across single-dialect and cross-dialect settings with initial results and story continuation, we analyze AI-generated content for coherence, creativity, and cultural relevance. Native Bengali speakers provide qualitative feedback, highlighting key challenges such as dialectal fidelity and narrative richness. Our findings emphasize the need for culturally adaptive NLP models to enhance storytelling in low-resource languages.},
    booktitle = {Companion Publication of the 2025 Conference on Computer-Supported Cooperative Work and Social Computing},
    pages = {214–218},
    numpages = {5},
    keywords = {Cultural bias, Large language models, Bengali language, Dialectal bias, Human-centered NLP, LLM auditing, Fairness and inclusion},
    location = {
    },
    series = {CSCW Companion '25}
    }
  28. LegalMind: An Intelligent Solution for Legal Document Analysis with User-Centric UI and AI-Driven Capabilities in Mobile Devices
    Azmine Toushik Wasi, Mst Rafia Islam, Abdur Rahman, Tawfia Yeasmin, Amit Agarwal, Hitesh Laxmichand Patel, TH Rafi, Dong-Kyu Chae
    CSCW 2025 Posters (A) â–Ș [PDF]

    Navigating legal documents is challenging due to their complexity, jargon, and interconnected entities like names, dates, and provisions, leading to inefficiencies and critical oversights. With the growing reliance on mobile devices, there is an increasing demand for tools that enable efficient and accessible legal document analysis on-the-go. To address these issues, we introduce LegalMind, a mobile app that revolutionizes legal document analysis by integrating key features: Automatic Entity Mapping for quick identification of essential details, Intelligent Question Answering for context-aware responses, and Multi-document Analysis for comprehensive comparison. Designed with a user-centric UI, LegalMind streamlines information retrieval, reduces manual effort, and enhances decision-making for legal professionals. The intuitive UI allows users to easily navigate complex legal content on mobile devices. Our user study confirmed that LegalMind improves efficiency and accessibility, thus showing its ability to transform legal workflows by bridging the gap between complex legal challenges and AI-driven solutions while maintaining a seamless, accessible UI.
    @inproceedings{10.1145/3715070.3749262,
    author = {Wasi, Azmine Toushik and Islam, Mst Rafia and Rahman, Abdur and Yeasmin, Tawfia and Agarwal, Amit and Patel, Hitesh Laxmichand and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {LegalMind: An Intelligent Solution for Legal Document Analysis with User-Centric UI and AI-Driven Capabilities in Mobile Devices},
    year = {2025},
    isbn = {9798400714801},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3715070.3749262},
    doi = {10.1145/3715070.3749262},
    abstract = {Navigating legal documents is challenging due to their complexity, jargon, and interconnected entities like names, dates, and provisions, leading to inefficiencies and critical oversights. With the growing reliance on mobile devices, there is an increasing demand for tools that enable efficient and accessible legal document analysis on-the-go. To address these issues, we introduce LegalMind, a mobile app that revolutionizes legal document analysis by integrating key features: Automatic Entity Mapping for quick identification of essential details, Intelligent Question Answering for context-aware responses, and Multi-document Analysis for comprehensive comparison. Designed with a user-centric UI, LegalMind streamlines information retrieval, reduces manual effort, and enhances decision-making for legal professionals. The intuitive UI allows users to easily navigate complex legal content on mobile devices. Our user study confirmed that LegalMind improves efficiency and accessibility, thus showing its ability to transform legal workflows by bridging the gap between complex legal challenges and AI-driven solutions while maintaining a seamless, accessible UI.},
    booktitle = {Companion Publication of the 2025 Conference on Computer-Supported Cooperative Work and Social Computing},
    pages = {410–414},
    numpages = {5},
    keywords = {Legal document analysis, Entity mapping, Question answering, Mobile LegalTech, AI-assisted legal workflow},
    location = {
    },
    series = {CSCW Companion '25}
    }
  29. Knowledge Explorer: An Agentic AI Framework for Interactive, Personalized and Multilingual Learning Experience
    Wahid Faisal*, Azmine Toushik Wasi*, Drishti Sharma*, Mahfuz Ahmed Anik, TH Rafi, Dong-Kyu Chae
    CSCW 2025 Posters (A) â–Ș [PDF]

    Digital education tools have made learning more accessible, yet they often fall short in providing personalized guidance, multilingual support, and engaging pedagogical strategies. As global learners seek more adaptive and inclusive educational experiences, there is a growing need for systems that replicate the explanatory depth of expert tutors. Existing solutions typically lack fine-grained topic breakdowns, contextual grounding, and dynamic delivery methods. To address these gaps, we present Knowledge Explorer, a multi-agent AI system that delivers personalized, multilingual explanations using structured topic decomposition, retrieval-augmented generation (RAG), and storytelling. The system leverages LangGraph and LangChain to coordinate agents for subtopic division, factual retrieval, and narrative-based teaching, integrating Cohere’s SOTA embedding, rerank and LLM models alongside ChromaDB as the vector store. Our key contributions include a language-agnostic architecture for education, a storytelling-based delivery agent, and a RAG pipeline grounded in authoritative sources. Preliminary results show effective content generation in English, Hindi, and Spanish, underscoring the system’s potential as a scalable and globally adaptable educational companion. Overall, these capabilities position Knowledge Explorer as a promising tool for advancing learning equity, particularly in underserved and linguistically diverse communities.
    @inproceedings{10.1145/3715070.3749282,
    author = {Faisal, Wahid and Wasi, Azmine Toushik and Sharma, Drishti and Anik, Mahfuz Ahmed and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {Knowledge Explorer: An Agentic AI Framework for Interactive, Personalized and Multilingual Learning Experience},
    year = {2025},
    isbn = {9798400714801},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3715070.3749282},
    doi = {10.1145/3715070.3749282},
    abstract = {Digital education tools have made learning more accessible, yet they often fall short in providing personalized guidance, multilingual support, and engaging pedagogical strategies. As global learners seek more adaptive and inclusive educational experiences, there is a growing need for systems that replicate the explanatory depth of expert tutors. Existing solutions typically lack fine-grained topic breakdowns, contextual grounding, and dynamic delivery methods. To address these gaps, we present Knowledge Explorer, a multi-agent AI system that delivers personalized, multilingual explanations using structured topic decomposition, retrieval-augmented generation (RAG), and storytelling. The system leverages LangGraph and LangChain to coordinate agents for subtopic division, factual retrieval, and narrative-based teaching, integrating Cohere’s SOTA embedding, rerank and LLM models alongside ChromaDB as the vector store. Our key contributions include a language-agnostic architecture for education, a storytelling-based delivery agent, and a RAG pipeline grounded in authoritative sources. Preliminary results show effective content generation in English, Hindi, and Spanish, underscoring the system’s potential as a scalable and globally adaptable educational companion. Overall, these capabilities position Knowledge Explorer as a promising tool for advancing learning equity, particularly in underserved and linguistically diverse communities.},
    booktitle = {Companion Publication of the 2025 Conference on Computer-Supported Cooperative Work and Social Computing},
    pages = {520–524},
    numpages = {5},
    keywords = {Topic decomposition, Multilingual education, Retrieval-augmented generation, Story-based learning, Personalized learning paths, AI-powered tutoring},
    location = {
    },
    series = {CSCW Companion '25}
    }
  30. Dialectal Bias in Bengali: An Evaluation of Multilingual Large Language Models Across Cultural Variations
    Azmine Toushik Wasi, Raima Islam, Mst Rafia Islam, Farig Sadeque, TH Rafi, Dong-Kyu Chae
    WWW 2025 (TheWebConf) (A*) (Short Paper) â–Ș [PDF]

    Large Language Models (LLMs) have transformed human-centric AI applications on the Web, yet they often exhibit stereotypes and biases, especially in sensitive contexts like cultural differences in low-resource languages such as Bengali. In this work, we investigate cultural bias in LLMs by evaluating their performance in Bengali cultural dialects of Hindu and Muslim majority. We evaluated widely used Web-enabled models, including ChatGPT, Gemini, and Microsoft Copilot, using a curated data set to analyze their handling of culturally specific terms and approaches to mitigating social biases. By addressing bias in language technologies that underpin the modern Web, our study contributes to advancing human-centered NLP and LLM auditing. Through a detailed exploration of bias causes and evaluation methods, our goal is to promote fairness and inclusion for more than 300 million Bengali speakers in the evolving ecosystem of the Web.
    @inproceedings{10.1145/3701716.3715468,
    author = {Wasi, Azmine Toushik and Islam, Raima and Islam, Mst Rafia and Sadeque, Farig and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {Dialectal Bias in Bengali: An Evaluation of Multilingual Large Language Models Across Cultural Variations},
    year = {2025},
    isbn = {9798400713316},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3701716.3715468},
    doi = {10.1145/3701716.3715468},
    abstract = {Large Language Models (LLMs) have transformed human-centric AI applications on the Web, yet they often exhibit stereotypes and biases, especially in sensitive contexts like cultural differences in low-resource languages such as Bengali. In this work, we investigate cultural bias in LLMs by evaluating their performance in Bengali cultural dialects of Hindu and Muslim majority. We evaluated widely used Web-enabled models, including ChatGPT, Gemini, and Microsoft Copilot, using a curated data set to analyze their handling of culturally specific terms and approaches to mitigating social biases. By addressing bias in language technologies that underpin the modern Web, our study contributes to advancing human-centered NLP and LLM auditing. Through a detailed exploration of bias causes and evaluation methods, our goal is to promote fairness and inclusion for more than 300 million Bengali speakers in the evolving ecosystem of the Web.},
    booktitle = {Companion Proceedings of the ACM on Web Conference 2025},
    pages = {1380–1384},
    numpages = {5},
    keywords = {Bengali language, LLM auditing, cultural bias, dialectal bias, fairness and inclusion, human-centered NLP, large language models},
    location = {Sydney NSW, Australia},
    series = {WWW '25}
    }
  31. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
    Angelika Romanou, Negar Foroutan, Anna Sotnikova,... , Azmine Toushik Wasi, ... , Marzieh Fadaee, Sara Hooker, Antoine Bosselut
    ICLR 2025 (Spotlight, Top 5.1%) (A*) â–Ș â–Ș [OpenReview]

    The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (i.e., multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
    @inproceedings{
    romanou2025include,
    title={{INCLUDE}: Evaluating Multilingual Language Understanding with Regional Knowledge},
    author={Angelika Romanou and Negar Foroutan and Anna Sotnikova and Sree Harsha Nelaturu and Shivalika Singh and Rishabh Maheshwary and Micol Altomare and Zeming Chen and Mohamed A. Haggag and Snegha A and Alfonso Amayuelas and Azril Hafizi Amirudin and Danylo Boiko and Michael Chang and Jenny Chim and Gal Cohen and Aditya Kumar Dalmia and Abraham Diress and Sharad Duwal and Daniil Dzenhaliou and Daniel Fernando Erazo Florez and Fabian Farestam and Joseph Marvin Imperial and Shayekh Bin Islam and Perttu Isotalo and Maral Jabbarishiviari and B{\"o}rje F. Karlsson and Eldar Khalilov and Christopher Klamm and Fajri Koto and Dominik Krzemi{\'n}ski and Gabriel Adriano de Melo and Syrielle Montariol and Yiyang Nan and Joel Niklaus and Jekaterina Novikova and Johan Samir Obando Ceron and Debjit Paul and Esther Ploeger and Jebish Purbey and Swati Rajwal and Selvan Sunitha Ravi and Sara Rydell and Roshan Santhosh and Drishti Sharma and Marjana Prifti Skenduli and Arshia Soltani Moakhar and Bardia soltani moakhar and Ayush Kumar Tarun and Azmine Toushik Wasi and Thenuka Ovin Weerasinghe and Serhan Yilmaz and Mike Zhang and Imanol Schlag and Marzieh Fadaee and Sara Hooker and Antoine Bosselut},
    booktitle={The Thirteenth International Conference on Learning Representations},
    year={2025},
    url={https://openreview.net/forum?id=k3gCieTXeY}
    }
  32. Gaussian Regularization in Neural Graph Learning
    Azmine Toushik Wasi, TH Rafi, Dong-Kyu Chae
    DASFAA 2025 (Full Paper, Short Oral, Top 10%) (B) â–Ș [PDF]

    Gaussian processes (GP), known for simplicity and flexibility, and the recent emergence of graph neural networks (GNNs) present promising avenues for semi-supervised learning on graph-structured data and beyond. Despite notable advancements in GNNs with a focus on neighborhood information, there are gaps in effectively integrating probability distributions and inter-entity relationships, as GNNs focus heavily on neighbors. Integrating probability distributions is essential to tackle noise and uncertainties. To address this issue, we present Gaussian Regularization in neural graph learning, which effectively incorporates Gaussian information to capture latent probabilistic distribution attributes from node embeddings based on the Gaussian Process, thereby boosting predictive capabilities. The process involves using a graph encoder to preprocess data, which is then encoded and used for effective aggregation and transformation of neighboring nodes. Gaussian Regularization works alongside any existing graph encoder model by encoding node representations into a Gaussian space to capture unique features. Logits are generated from this space, along with another set from an MLP, and then merged for final predictions. Extensive experiments using real-world benchmark datasets show that our approach outperforms several state-of-the-art GNN models and has a significant positive impact on off-the-shelf graph encoders and kernels, demonstrating its effectiveness and flexibility.
    @inbook{Wasi2026,
      title = {Gaussian Regularization in Neural Graph Learning},
      ISBN = {9789819538270},
      ISSN = {1611-3349},
      url = {http://dx.doi.org/10.1007/978-981-95-3827-0_11},
      DOI = {10.1007/978-981-95-3827-0_11},
      booktitle = {Database Systems for Advanced Applications},
      publisher = {Springer Nature Singapore},
      author = {Wasi,  Amzine Toushik and Rafi,  Taki Hasan and Chae,  Dong-Kyu},
      year = {2026},
      pages = {166–181}
    }
  33. Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models
    Ferdinand Lenchau Bit, Iman Khaleda binti Zamri, Azmine Toushik Wasi$, TH Rafi, Dong-Kyu Chae
    DASFAA 2025 (Short Paper Paper, Poster) (B) â–Ș [PDF]

    Language models (LMs) are essential for natural language processing (NLP) tasks, but they often exhibit biases due to inadequate or imbalanced training data, particularly in multilingual settings. These biases can lead to challenges in modeling linguistically similar low-resource languages, such as Malay and Indonesian, where mutual intelligibility complicates language differentiation. Addressing these biases is critical for enhancing the performance and fairness of NLP tools for underrepresented languages. Current LMs struggle with consistency in such scenarios, often leading to language mixing or poor prediction accuracy due to insufficient data capturing subtle linguistic differences. To tackle this, we curate a novel dataset of Malay sentences infused with Indonesian intrusions by simulating mixed-language sentences through filtering and refinement. We then fine-tune a RoBERTa model on this dataset. Empirically, this model exhibits significant improvements in word-level accuracy and language consistency compared to baseline models, indicating its ability to mitigate biases effectively. We believe that our work offers a pathway to address linguistic gaps and fosters the development of more accurate and equitable NLP tools for low-resource languages in multilingual environments.
    @inproceedings{10.1007/978-981-95-3827-0_23,
    author = {Bit, Ferdinand Lenchau and binti Zamri, Iman Khaleda and Wasi, Amzine Toushik and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {Mitigating Linguistic Bias Between Malay and Indonesian Languages Using Masked Language Models},
    year = {2026},
    isbn = {978-981-95-3826-3},
    publisher = {Springer-Verlag},
    address = {Berlin, Heidelberg},
    url = {https://doi.org/10.1007/978-981-95-3827-0_23},
    doi = {10.1007/978-981-95-3827-0_23},
    abstract = {Language models (LMs) are essential for natural language processing (NLP) tasks, but they often exhibit biases due to inadequate or imbalanced training data, particularly in multilingual settings. These biases can lead to challenges in modeling linguistically similar low-resource languages, such as Malay and Indonesian, where mutual intelligibility complicates language differentiation. Addressing these biases is critical for enhancing the performance and fairness of NLP tools for underrepresented languages. Current LMs struggle with consistency in such scenarios, often leading to language mixing or poor prediction accuracy due to insufficient data capturing subtle linguistic differences. To tackle this, we curate a novel dataset of Malay sentences infused with Indonesian intrusions by simulating mixed-language sentences through filtering and refinement. We then fine-tune a RoBERTa model on this dataset. Empirically, this model exhibits significant improvements in word-level accuracy and language consistency compared to baseline models, indicating its ability to mitigate biases effectively. We believe that our work offers a pathway to address linguistic gaps and fosters the development of more accurate and equitable NLP tools for low-resource languages in multilingual environments.},
    booktitle = {Database Systems for Advanced Applications: 30th International Conference, DASFAA 2025, Singapore, Singapore, May 26–29, 2025, Proceedings, Part I},
    pages = {328–338},
    numpages = {11},
    keywords = {Bias in NLP, Malay and Indonesian language models, NLP for underrepresented languages},
    location = {Singapore, Singapore}
    }
  34. Physics-Informed Neural Networks for Clinical Time-Series Forecasting with Clinical Utility
    Azmine Toushik Wasi, Mahfuz Ahmed Anik, Abdur Rahman, Md Iqramul Hoque, Wahid Faisal, and A. M. M. Mukaddes
    ICCIT 2025 â–Ș [PDF]

    Clinical time-series forecasting is essential for anticipating patient deterioration and supporting proactive interventions, yet traditional machine learning approaches often operate as black boxes, limiting trust and adoption in critical care. To address this challenge, we implement a Physics-Informed Neural Network (PINN) that integrates physiological constraints, such as fluid balance and pharmacokinetics, directly into the learning process, ensuring predictions remain consistent with known medical principles. The study focuses on forecasting laboratory and vital sign trends while simultaneously predicting deterioration risk (ICU stay or in-hospital death) from a dataset containing demographics, preoperative labs, and intraoperative measurements. By embedding physical laws into the loss function, the model enhances interpretability, generalization, and plausibility, offering a bridge between data-driven predictions and physiological reasoning. Results show strong regression performance (RMSE = 11.75, MAE = 8.46) and competitive classification metrics (Accuracy = 0.84, AUC=0.82), though recall highlights the challenge of capturing all high-risk cases. These findings demonstrate that PINNs not only improve predictive robustness but also provide clinically meaningful insights, marking a step toward trustworthy, physically grounded AI for healthcare.
    @INPROCEEDINGS{11489535,
      author={Wasi, Azmine Toushik and Anik, Mahfuz Ahmed and Rahman, Abdur and Hoque, Md Iqramul and Faisal, Wahid and Mukaddes, A. M. M.},
      booktitle={2025 28th International Conference on Computer and Information Technology (ICCIT)}, 
      title={Physics-Informed Neural Networks for Clinical Time-Series Forecasting with Clinical Utility}, 
      year={2025},
      volume={},
      number={},
      pages={1028-1033},
      keywords={Protocols;Network architecture;HTTP;Communication systems;LoRa;Radio access networks;Regional area networks;Data communication;Electronic messaging;Internet;Clinical time-series forecasting;Physics-Informed Neural Networks (PINNs);Physiological constraints;Patient deterioration prediction;Interpretability in healthcare AI},
      doi={10.1109/ICCIT68739.2025.11489535}}
    
  35. Sleep Quality Using the Pittsburgh Sleep Quality Index: A Machine Learning Approach
    Abdur Rahman, Mahfuz Ahmed Anik, Shahoriar Muttaki Utshaw, Azmine Toushik Wasi$, and A. M. M. Mukaddes
    ICCIT 2025 â–Ș [PDF]

    Sleep quality is a critical determinant of physical health, cognitive performance, and emotional well-being, yet poor sleep has become alarmingly prevalent among university students, particularly in Bangladesh. Existing research often relies on linear models and generalized recommendations, overlooking the complex interplay of behavioral, psychosocial, and environmental determinants. This study addresses that gap by applying a multifactorial machine learning framework to predict sleep quality using the Pittsburgh Sleep Quality Index (PSQI) alongside a comprehensive set of demographic, lifestyle, psychosocial, and environmental variables. Among the models tested, advanced ensemble methods demonstrated superior predictive capability, with XGBoost achieving the highest performance (F1=0.9544, Accuracy = 0.9366). Feature importance analysis using SHAP shows that health issues such as stress from financial pressures, frequent headaches and neck pain, environmental factors like lighting, ventilation, and quietness, as well as bedtime habits including screen use, posture, and caffeine intake, strongly affect sleep quality. The results suggest that reducing stress, improving sleep environments, adopting healthier routines, and maintaining a proper BMI can significantly enhance the well-being of students and the young generation.
    @INPROCEEDINGS{11491297,
      author={Rahman, Abdur and Anik, Mahfuz Ahmed and Utshaw, Shahoriar Muttaki and Wasi, Azmine Toushik and Mukaddes, A. M. M.},
      booktitle={2025 28th International Conference on Computer and Information Technology (ICCIT)}, 
      title={Predicting Sleep Quality Using the Pittsburgh Sleep Quality Index: A Machine Learning Approach}, 
      year={2025},
      volume={},
      number={},
      pages={1244-1249},
      keywords={Radio broadcasting;Frequency modulation;HTTP;Internet;Protocols;Communication systems;Computer networks;Modulation;Radio broadcasting;Frequency modulation;Sleep Quality;Pittsburgh Sleep Quality Index (PSQI);Machine Learning;Predictive Modelling},
      doi={10.1109/ICCIT68739.2025.11491297}}
    
  36. GReFEL: Geometry-Aware Reliable Facial Expression Learning under Bias and Imbalanced Data Distribution
    Azmine Toushik Wasi*, TH Rafi*, Raima Islam, Karlo Ć erbetar, Dong-Kyu Chae
    ACCV 2024 (B) â–Ș [CVF] â–Ș [LNCS] â–Ș [arXiv]

    Reliable facial expression learning (FEL) involves the effective learning of distinctive facial expression characteristics for more reliable, unbiased and accurate predictions in real-life settings. However, current systems struggle with FEL tasks because of the variance in people’s facial expressions due to their unique facial structures, movements, tones, and demographics. Biased and imbalanced datasets compound this challenge, leading to wrong and biased prediction labels. To tackle these, we introduce GReFEL, leveraging Vision Transformers and a facial geometry-aware anchor-based reliability balancing module to combat imbalanced data distributions, bias, and uncertainty in facial expression learning. Integrating local and global data with anchors that learn different facial data points and structural features, our approach adjusts biased and mislabeled emotions caused by intra-class disparity, inter-class similarity, and scale sensitivity, resulting in comprehensive, accurate, and reliable facial expression predictions. Our model outperforms current state-of-the-art methodologies, as demonstrated by extensive experiments on various datasets.
    @inbook{Wasi2024,
      title = {GReFEL: Geometry-Aware Reliable Facial Expression Learning Under Bias and Imbalanced Data Distribution},
      ISBN = {9789819609116},
      ISSN = {1611-3349},
      url = {http://dx.doi.org/10.1007/978-981-96-0911-6_27},
      DOI = {10.1007/978-981-96-0911-6_27},
      booktitle = {Computer Vision – ACCV 2024},
      publisher = {Springer Nature Singapore},
      author = {Wasi,  Azmine Toushik and Rafi,  Taki Hasan and Islam,  Raima and S̆erbetar,  Karlo and Chae,  Dong-Kyu},
      year = {2024},
      month = Dec,
      pages = {465–482}
    }
  37. BanglaAutoKG: Automatic Bangla Knowledge Graph Construction with Semantic Neural Graph Filtering
    Azmine Toushik Wasi, TH Rafi, Raima Islam, Dong-Kyu Chae
    COLING 2024 (B) â–Ș [ACL Anthology] â–Ș [arXiv] â–Ș [GitHub]

    Knowledge Graphs (KGs) have proven essential in information processing and reasoning applications because they link related entities and give context-rich information, supporting efficient information retrieval and knowledge discovery; presenting information flow in a very effective manner. Despite being widely used globally, Bangla is relatively underrepresented in KGs due to a lack of comprehensive datasets, encoders, NER (named entity recognition) models, POS (part-of-speech) taggers, and lemmatizers, hindering efficient information processing and reasoning applications in the language. Addressing the KG scarcity in Bengali, we propose BanglaAutoKG, a pioneering framework that is able to automatically construct Bengali KGs from any Bangla text. We utilize multilingual LLMs to understand various languages and correlate entities and relations universally. By employing a translation dictionary to identify English equivalents and extracting word features from pre-trained BERT models, we construct the foundational KG. To reduce noise and align word embeddings with our goal, we employ graph-based polynomial filters. Lastly, we implement a GNN-based semantic filter, which elevates contextual understanding and trims unnecessary edges, culminating in the formation of the definitive KG. Empirical findings and case studies demonstrate the universal effectiveness of our model, capable of autonomously constructing semantically enriched KGs from any text. Data and code are available here: https://github.com/azminewasi/BanglaAutoKG
    @inproceedings{wasi-etal-2024-banglaautokg,
        title = "{B}angla{A}uto{KG}: Automatic {B}angla Knowledge Graph Construction with Semantic Neural Graph Filtering",
        author = "Wasi, Azmine Toushik  and Rafi, Taki Hasan  and Islam, Raima  and Chae, Dong-Kyu",
        booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
        month = may,
        year = "2024",
        address = "Torino, Italia",
        publisher = "ELRA and ICCL",
        url = "https://aclanthology.org/2024.lrec-main.189/",
        pages = "2100--2106",
    }
  38. Neural Control System for Continuous Glucose Monitoring and Maintenance
    Azmine Toushik Wasi
    ICLR 2024 (A*) Tiny Papers â–Ș [OpenReview] â–Ș [arXiv] â–Ș [GitHub]

    Precise glucose level monitoring is critical for people with diabetes to avoid serious complications. While there are several methods for continuous glucose level monitoring, research on maintenance devices is limited. To mitigate the gap, we provide a novel neural control system for continuous glucose monitoring and management that uses differential predictive control, NeuralCGMM. Our approach, led by a sophisticated neural policy and differentiable modeling, constantly adjusts insulin supply in real-time, thereby improving glucose level optimization in the body. This end-to-end method maximizes efficiency, providing personalized care and improved h
    @inproceedings{
    wasi2024neural,
    title={Neural Control System for Continuous Glucose Monitoring and Maintenance},
    author={Azmine Toushik Wasi},
    booktitle={The Second Tiny Papers Track at ICLR 2024},
    year={2024},
    url={https://openreview.net/forum?id=Te4P3Cn54g}
    }
  39. When SMILES have Language: Drug Classification using Text Classification Methods on Drug SMILES Strings
    Azmine Toushik Wasi, Karlo Serbetar, Raima Islam, TH Rafi, Dong-Kyu Chae
    ICLR 2024 (A*) Tiny Papers â–Ș [OpenReview] â–Ș [arXiv] â–Ș [GitHub]

    Complex chemical structures, like drugs, are usually defined by SMILES strings as a sequence of molecules and bonds. These SMILES strings are used in different complex machine learning-based drug-related research and representation works. Escaping from complex representation, in this work, we pose a single question: What if we treat drug SMILES as conventional sentences and engage in text classification for drug classification? Our experiments affirm the possibility with very competitive scores. The study explores the notion of viewing each atom and bond as sentence components, employing basic NLP methods to categorize drug types, proving that complex problems can also be solved with simpler perspectives. The data and code are available here: https://github.com/azminewasi/Drug-Classification-NLP.
    @inproceedings{
    wasi2024when,
    title={When {SMILES} have Language: Drug Classification using Text Classification Methods on Drug {SMILES} Strings},
    author={Azmine Toushik Wasi and {\v{S}}erbetar Karlo and Raima Islam and Taki Hasan Rafi and Dong-Kyu Chae},
    booktitle={The Second Tiny Papers Track at ICLR 2024},
    year={2024},
    url={https://openreview.net/forum?id=VUYCyH8fCw}
    }
  40. DiaFrame: A Framework for Understanding Bengali Dialects in Human-AI Collaborative Creative Writing Spaces
    Azmine Toushik Wasi, TH Rafi, Dong-Kyu Chae
    CSCW 2024 Posters (A) â–Ș [DOI]

    Preserving linguistic dialects is paramount for preserving cultural identity, fostering diversity, and maintaining cultural prestige. In the context of Bengali, the language exhibits religious distinctions through its two major dialects: the West Bengal (Hindu majority) and Bangladesh (Muslim majority) dialects, deeply rooted in the historical evolution of Bangla. Understanding and respecting these dialect-based nuances is crucial for Human-AI collaborative tools to generate culturally appropriate content and effectively assist human writers without limiting their creativity. To address these challenges, we introduce DiaFrame, a comprehensive framework tailored for understanding and tuning dialects in human-AI collaborative creative writing spaces, focusing on Bengali. It integrates advanced components such as active learning, real-time processing of human feedback, memory management, and contextual understanding to enhance user experience and ensure culturally sensitive content generation. Our contributions include providing a holistic end-to-end solution that enables the AI model to actively learn, adapt, and generate content aligned with user preferences and dialectical variations, ultimately fostering effective collaboration between humans and AI in creative writing endeavors.
    @inproceedings{10.1145/3678884.3681862,
    author = {Wasi, Azmine Toushik and Rafi, Taki Hasan and Chae, Dong-Kyu},
    title = {DiaFrame: A Framework for Understanding Bengali Dialects in Human-AI Collaborative Creative Writing Spaces},
    year = {2024},
    isbn = {9798400711145},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3678884.3681862},
    doi = {10.1145/3678884.3681862},
    abstract = {Preserving linguistic dialects is paramount for preserving cultural identity, fostering diversity, and maintaining cultural prestige. In the context of Bengali, the language exhibits religious distinctions through its two major dialects: the West Bengal (Hindu majority) and Bangladesh (Muslim majority) dialects, deeply rooted in the historical evolution of Bangla. Understanding and respecting these dialect-based nuances is crucial for Human-AI collaborative tools to generate culturally appropriate content and effectively assist human writers without limiting their creativity. To address these challenges, we introduce DiaFrame, a comprehensive framework tailored for understanding and tuning dialects in human-AI collaborative creative writing spaces, focusing on Bengali. It integrates advanced components such as active learning, real-time processing of human feedback, memory management, and contextual understanding to enhance user experience and ensure culturally sensitive content generation. Our contributions include providing a holistic end-to-end solution that enables the AI model to actively learn, adapt, and generate content aligned with user preferences and dialectical variations, ultimately fostering effective collaboration between humans and AI in creative writing endeavors.},
    booktitle = {Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing},
    pages = {268–274},
    numpages = {7},
    keywords = {bangla language, dialogues and discourse, human-ai collaborative writing spaces, languages and dialects, large language models},
    location = {San Jose, Costa Rica},
    series = {CSCW Companion '24}
    }


📰 Journal Papers


  1. CADGL: Context-Aware Deep Graph Learning for Predicting Drug-Drug Interactions
    Azmine Toushik Wasi, TH Rafi, Raima Islam, Serbetar Karlo, Dong-Kyu Chae
    IEEE/ACM Transactions on Computational Biology and Bioinformatics (TCBB) (Q1/Q2, IF: 4.5, CiteScore: 9.7) â–Ș [DOI] â–Ș [arXiv]

    Examining Drug-Drug Interactions (DDIs) is a pivotal element in the process of drug development. DDIs occur when one drug's properties are affected by the inclusion of other drugs. Detecting favorable DDIs has the potential to pave the way for creating and advancing innovative medications applicable in practical settings. However, existing DDI prediction models continue to face challenges related to generalization in extreme cases, robust feature extraction, and real-life application possibilities. We aim to address these challenges by leveraging the effectiveness of context-aware deep graph learning by introducing a novel framework named CADGL. Based on a customized variational graph autoencoder (VGAE), we capture critical structural and physio-chemical information using two context pre-processors for feature extraction from two different perspectives: local neighborhood and molecular context, in a heterogeneous graphical structure. Our customized VGAE consists of a graph encoder, a latent information encoder, and an MLP decoder. CADGL surpasses other state-of-the-art DDI prediction models, excelling in predicting clinically valuable novel DDIs, supported by rigorous case studies.
    @ARTICLE{11441444,
      author={Wasi, Azmine Toushik and Rafi, Taki Hasan and Islam, Raima and Karlo, Ć erbetar and Chae, Dong-Kyu},
      journal={IEEE Transactions on Computational Biology and Bioinformatics}, 
      title={CADGL: Context-Aware Deep Graph Learning for Predicting Drug-Drug Interactions}, 
      year={2026},
      volume={23},
      number={3},
      pages={1265-1276},
      keywords={Drugs;Feature extraction;Predictive models;Decoding;Autoencoders;Computational modeling;Vectors;Shape;Chemicals;Faces;Graph neural networks;computational drug discovery;drug-drug interactions;molecular interaction networks;pharmaceutical informatics;context-aware learning},
      doi={10.1109/TCBBIO.2026.3675142}}
    
  2. A Computational Community Blind Challenge on Pan-Coronavirus Drug Discovery Data
    Hugo MacDermott-Opeskin, Cas Wognum, ..., Azmine Toushik Wasi, ..., James Fraser, John D. Chodera
    Journal of Chemical Information and Modeling (JCIM) (Q1, IF 5.3) â–Ș [ACS] [chemrXiv]

    Computational blind challenges offer critical, unbiased opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the past decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in drug discovery, using lead optimization data from the AI-driven Structure-enabled Antiviral Platform Discovery Consortium’s pan-coronavirus antiviral discovery program, in partnership with Polaris and the OpenADMET project. This collaborative initiative invited global participants from both academia and industry to develop and apply computational methods to predict the biochemical potency and crystallographic ligand poses of small molecules against key coronavirus targets, Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) and Middle East Respiratory Syndrome Coronavirus (MERS-CoV) main protease (Mpro), as well as multiple ADMET assay end points, using previously undisclosed comprehensive experimental drug discovery data sets as benchmarks. By evaluating submissions across multiple tasks and compounds, we established performance leaderboards and conducted meta-analyses to assess methodological strengths, common pitfalls, and areas for improvement. This analysis provides a foundation for best practices in real-world machine learning evaluation, grounded in community-driven benchmarking. We also highlight how next-generation platforms, such as Polaris, enable rigorous challenge design, embedded evaluation frameworks, and broad community engagement. This paper reports the collective findings of the challenge, offering a high-level overview of the data, evaluation infrastructure, and top-performing strategies. We further provide context and support for the accompanying papers authored by the challenge participants in this special issue, which explore individual approaches in greater depth. Together, these contributions aim to advance reproducible, trustworthy, and high-impact computational methods in drug discovery, and to explore best practices and pitfalls in future blind challenge design and execution, including planned initiatives for the OpenADMET project.
    @article{MacDermottOpeskin2026,
      title = {A Computational Community Blind Challenge on Pan-Coronavirus Drug Discovery Data},
      volume = {66},
      ISSN = {1549-960X},
      url = {http://dx.doi.org/10.1021/acs.jcim.5c02106},
      DOI = {10.1021/acs.jcim.5c02106},
      number = {6},
      journal = {Journal of Chemical Information and Modeling},
      publisher = {American Chemical Society (ACS)},
      author = {MacDermott-Opeskin,  Hugo and Scheen,  Jenke and Wognum,  Cas and Horton,  Joshua T. and West,  Devany and Payne,  Alexander Matthew and Castellanos,  Maria A. and Colby,  Sean and Griffen,  Edward and Cousins,  David and Stacey,  Jessica and Reid,  Lauren and Aschenbrenner,  Jasmin Cara and Fearon,  Daren and Balcomb,  Blake and Marples,  Peter and Tomlinson,  Charles W. E. and Lithgo,  Ryan and Godoy,  Andre S. and Winokan,  Max and Barr,  Haim and Lahav,  Noa and Lavi,  Michael and Duberstein,  Shirley and Cohen,  Galit and Fate,  Gwendolyn and Lefker,  Bruce and Robinson,  Ralph and Szommer,  Tamas and Lynch,  Nick and Minh,  David D. L. and La,  Van Ngoc Thuy and Kang,  Lulu and Huddleston,  Kate and Renslow,  Ryan and Tollefson,  Mallory and Walters,  W. Patrick and Xu,  Cynthia and Hsu,  Jonny and St-Laurent,  Julien and Etsmoberg,  Honore and Zhu,  Lu and Quirke,  Andrew and Abdul Haleem,  Mohamed Iliyas and Alibay,  Irfan and Baid,  Gunjan and Birnbaum,  Benjamin and Bishop,  Kevin P. and Bohorquez,  Hugo and Bose,  Ashmita and Brown,  C. J. and Burns,  Jackson and Cai,  Lianjin and Cedeno,  Ruel and de Cesco,  Stephane and Chupakhin,  Vladimir and Clark,  Finlay and Cole,  Daniel J. and Corbi-Verge,  Carles and Danial,  Muhammad and Davi,  Alec and Dehaen,  Wim and Doering,  Niklas Piet and Dougha,  Alexis and Dréanic,  Marie-Pierre and Eakin,  Bryce and Ehrlich,  Anatol and Elijosius,  Rokas and F\"{u}l\"{o}p,  Jozef and Gitter,  Anthony and Goossens,  Kenneth and Gu,  Yaowen and Head-Gordon,  Teresa and Hoffer,  Laurent and Hofmans,  Johan and Jiang,  Ellena and Kaminow,  Benjamin and Khosravi,  Sina and Khoualdi,  Asma Feriel and Lenselink,  Eelke Bart and Liu,  Zhirong and Liu,  Yue and Liu,  Sijie and Ma,  Yizhou and Maher,  Patrick and Mayer,  Imke and Mendez-Lucio,  Oscar and Mey,  Antonia S. J. S. and Michel,  Julien and Montanari,  Floriane and Niu,  Taoyu and Ogino,  Ryusei and Palaniappan,  Ashok and Pan,  Xiaolin and Patnaik,  Auro and Pham,  Long-Hung and Pinto,  Luis and Purnomo,  Justin and Rich,  Alex and Schaaf,  Lars and Schran,  Christoph and Singh,  Rajeev Kumar and Srilakshmi,  Mounika and Srivastava,  Satya Pratik and Sun,  Kunyang and Sun,  Zhaoxi and Talagayev,  Valerij and Thirukonda Subramanian Balakrishnan,  Balamurugan and Titus,  Ida and Tkatchenko,  Alexandre and Treyde,  Wojtek and Tricarico,  Giovanni and Tripp,  Austin and Vithayapalert,  Nopsinth and Wang,  Yingze and Wasi,  Azmine Toushik and Wedig,  Steffen and Wolber,  Gerhard and Xu,  Bofei and Zhou,  Weijun and von Delft,  Frank and Lee,  Alpha and Kirkegaard,  Karla and Sj\"{o},  Peter and Fraser,  James S. and Chodera,  John D.},
      year = {2026},
      month = Feb,
      pages = {3129–3149}
    }
  3. GRAPHINE: Enhancing Spatiotemporal Supply Chain Forecasting Using Virtual Node-Augmented Graph Diffusion for Improved Fuel Efficiency
    A. S. A.*, Azmine Toushik Wasi*, Mahfuz Ahmed Anik, MD Shafikul Islam, and Mohamed Kamel Hadj-Kali
    International Journal of Production Research (IJPR) (Q1) [DOI]

    Spatiotemporal forecasting in supply chain networks demands modelling complex spatial dependencies and nonlinear temporal dynamics. Traditional models often fail to capture the heterogeneity, structural sparsity, and long-term dependencies of real-world supply chains. To address this, we propose GRAPHINE, a Virtual Node Diffusion-Convolutional Recurrent Neural Network tailored for supply chain forecasting. GRAPHINE leverages diffusion-based learning, using bidirectional random walks to model spatial relations and a recurrent encoder-decoder with scheduled sampling to capture temporal patterns. It introduces virtual nodes to aggregate global context and applies a learnable gating mechanism that allows each node to regulate global influence based on local features, preventing oversmoothing and preserving specificity. Evaluated on the SCG dataset for demand and inventory forecasting, GRAPHINE achieves reductions of 38.71% in MSE and of 21.71% in RMSE over state-of-the-art baselines, which may imply potential reductions in fuel use on the order of 5.3% in scenario-specific settings under commonly used elasticity assumptions in logistics planning. GRAPHINE thus establishes a new decision-aid framework for applying advanced spatiotemporal graph neural architectures to complex production and supply chain problems, with tangible implications for production planning, operational cost efficiency, and environmental sustainability.
    @article{Alshehri2026,
      title = {GRAPHINE: enhancing spatiotemporal supply chain forecasting using virtual node-augmented graph diffusion for improved fuel efficiency},
      ISSN = {1366-588X},
      url = {http://dx.doi.org/10.1080/00207543.2026.2685179},
      DOI = {10.1080/00207543.2026.2685179},
      journal = {International Journal of Production Research},
      publisher = {Informa UK Limited},
      author = {Alshehri,  Abdulelah S. and Wasi,  Azmine Toushik and Anik,  Mahfuz Ahmed and Islam,  MD Shafikul and Hadj-Kali,  Mohamed Kamel},
      year = {2026},
      month = June,
      pages = {1–26}
    }
  4. AI Collaboration and the Decentering of Human Creativity
    Azmine Toushik Wasi, Sadia Tasneem Meem
    ACM AI Letters [DOI] â–Ș

    Western philosophical traditions have long framed creativity as an exclusively human capacity, reinforcing anthropocentric assumptions about authorship and artistic agency. The rise of generative AI challenges this view by introducing non-human systems capable of contributing meaningfully to creative production. This article examines how AI collaboration destabilizes traditional human-centered models of creativity. Drawing on posthumanist and new materialist perspectives, we argue that creativity should be understood as a distributed process emerging from interactions between human and non-human agents. This perspective calls for rethinking authorship, originality, and creative agency in an era of human–AI collaboration.
    @article{10.1145/3816258,
    author = {Wasi, Azmine Toushik and Meem, Sadia Tasnim},
    title = {AI Collaboration and the Decentering of Human Creativity},
    year = {2026},
    issue_date = {June 2026},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    volume = {1},
    number = {2},
    url = {https://doi.org/10.1145/3816258},
    doi = {10.1145/3816258},
    abstract = {Western philosophical traditions have long framed creativity as an exclusively human capacity, reinforcing anthropocentric assumptions about authorship and artistic agency. The rise of generative AI challenges this view by introducing non-human systems capable of contributing meaningfully to creative production. This article examines how AI collaboration destabilizes traditional human-centered models of creativity. Drawing on posthumanist and new materialist perspectives, we argue that creativity should be understood as a distributed process emerging from interactions between human and non-human agents. This perspective calls for rethinking authorship, originality, and creative agency in an era of human–AI collaboration.},
    journal = {ACM AI Lett.},
    month = jun,
    articleno = {13},
    numpages = {6},
    keywords = {Posthumanism, Generative AI, Human–AI Collaboration, Creative Agency, Authorship, New Materialism, Computational Creativity, AI Art}
    }
  5. ST-ResGAT: Explainable Spatio-Temporal Graph Neural Network for Road Condition Prediction and Priority-Driven Maintenance
    Mohsin Mahmud Topu#, Azmine Toushik Wasi#, Mahfuz Ahmed Anik, MD Manjurul Ahsan
    Intelligent Transportation Infrastructure (Q1) â–Ș [DOI]

    Climate-vulnerable road networks demand a transition from reactive, fix-on-failure maintenance toward predictive anddecision-ready strategies. Addressing this need, we propose ST-ResGAT, a novel spatio-temporal residual graph attentionnetwork that integrates residual graph-attention encoding with GRU-based temporal aggregation to model and forecastpavement deterioration. The framework is explicitly designed for resource-constrained environments and directly maps con-tinuous pavement condition index predictions to American Society for Testing and Materials (ASTM)-compliant maintenancepriorities. We evaluate the proposed approach on a real-world inspection dataset comprising 750 road segments in Sylhet,Bangladesh, collected between 2021 and 2024. Experimental results demonstrate that ST-ResGAT substantially outperformsconventional non-spatial machine learning baselines, achieving high predictive accuracy (R2 = 0.93, RMSE = 2.72). Ablationanalysis further reveals the critical role of spatial topology, confirming that pavement degradation propagates throughnetwork connectivity. To enhance interpretability, we incorporate GNNExplainer, showing that the model’s learned decisionpatterns align with established engineering principles. In addition, we assess classification reliability in practical deploymentscenarios, achieving 85.5% exact ASTM class agreement and 100% adjacent-class containment, thereby ensuring boundedand engineer-safe predictions. These results highlight ST-ResGAT as a practical, explainable, and robust solution for intelligentinfrastructure management in high-risk, resource-limited settings.
    @article{Topu2026,
      title = {ST-ResGAT: explainable spatio-temporal graph neural network for road condition prediction and priority-driven maintenance},
      volume = {5},
      ISSN = {2752-9991},
      url = {http://dx.doi.org/10.1093/iti/liag006},
      DOI = {10.1093/iti/liag006},
      journal = {Intelligent Transportation Infrastructure},
      publisher = {Oxford University Press (OUP)},
      author = {Topu,  Mohsin Mahmud and Wasi,  Azmine Toushik and Anik,  Mahfuz Ahmed and Ahsan,  M D Manjurul},
      year = {2026}
    }
  6. Digital Twin Enabled Additive Manufacturing: A Comprehensive Review of Architectures, Integration Layers, and Operational Maturity
    Mahfuz Ahmed Anik, Abdur Rahman, MD Shafikul Islam, Md Isfar Khan, Md Manjurul Ahsan, Azmine Toushik Wasi$#
    Digital Twins and Applications â–Ș [DOI]

    Digital twin-enabled additive manufacturing (DT-AM) represents a critical advancement in industrial intelligence, yet its current fragmented implementations necessitate a comprehensive systematic analysis. This review rigorously examines DT-AM architectures, integration frameworks, and pathways towards achieving autonomous and scalable production environments. Employing a structured multidimensional taxonomy, DT-AM systems are categorised based on functional scope (component, asset, system, process twins), integration depth (digital model, digital shadow, digital twin) and operational sophistication (ranging from descriptive to fully autonomous twins). The synthesis highlights a pronounced emphasis in existing literature on simulation and control functionalities, with notable gaps identified in validation frameworks and intelligence-driven decision-making mechanisms. Major challenges, including data heterogeneity, computational scalability and inadequate validation strategies, currently hinder seamless interoperability and broader industrial adoption. The analysis further identifies promising technological solutions such as agentic artificial intelligence, secure digital thread infrastructures, hybrid cloud-edge computing and human-centric augmented and virtual reality interfaces. Crucially, the paper underscores the imperative of software standardisation as foundational to the progression of DT-AM systems. Five strategic research trajectories are proposed to systematically bridge existing technological and operational gaps, fostering resilient, interoperable and human-centric cyber-physical manufacturing ecosystems. Ultimately, this comprehensive review establishes a structured foundation for advancing DT-AM, guiding future scholarly and industrial efforts towards cohesive, intelligent and scalable production systems.
    @article{Anik2026,
      title = {Digital Twin‐Enabled Additive Manufacturing: A Comprehensive Review of Architectures,  Integration Layers and Operational Maturity},
      volume = {3},
      ISSN = {2995-2182},
      url = {http://dx.doi.org/10.1049/dgt2.70039},
      DOI = {10.1049/dgt2.70039},
      number = {1},
      journal = {Digital Twins and Applications},
      publisher = {Institution of Engineering and Technology (IET)},
      author = {Anik,  Mahfuz Ahmed and Rahman,  Abdur and Islam,  MD Shafikul and Khan,  Md Isfar and Ahsan,  Md Manjurul and Wasi,  Azmine Toushik},
      year = {2026},
      month = Jan 
    }
  7. Beyond Local Receptive Fields: Vision Transformers for Real-Time Surface Defect Detection in FDM
    MD Shafikul Islam, Mahathir Mohammad Bappy, Saifur Rahman Tushar, Christian Zamiela, Azmine Toushik Wasi
    International Journal of Advanced Manufacturing Technology (Q1) [PDF]

    Ensuring real-time quality assurance in additive manufacturing (AM), particularly Fused Deposition Modeling (FDM), is essential due to its widespread adoption across industries driven by its process efficiency, design flexibility, and low material waste. However, conventional defect detection methods exhibit fundamental limitations in capturing spatially distributed and subtle surface anomalies. These limitations stem from their localized receptive fields and strong inductive biases, which restrict their ability to generalize to complex surface patterns. To address these challenges, this study proposes an explicitly optimized method based on Vision Transformers (ViTs) for real-time surface anomaly detection in the FDM process. We evaluate four surface conditions (normal, under-extrusion, over-extrusion, and void/empty regions) and generate explainability outputs on-demand to preserve real-time monitoring. Unlike CNNs, ViTs utilize global self-attention mechanisms, enabling them to capture long-range dependencies and subtle spatial variations across the printed surface. This methodological advantage allows for enhanced sensitivity to defect characteristics that conventional models often overlook. The method integrates depth maps derived from 2D laser scanning to construct high-fidelity surface topology representations, facilitating accurate classification of key defect types, including under-extrusion, over-extrusion, voids, and normal printing. To support practical deployment, we include post-hoc explanation modules based on attention visualization, gradient-based attribution (Integrated Gradients/ Saliency), and embedding-space projection (t-SNE and UMAP) to provide operator-facing evidence of regions and representation structure associated with each prediction. Experimental evaluation achieves a macro-F1 score of 0.877 and a macro-averaged mean AUC (mAUC) of 0.972, with an inference latency of 14.80–32.45 ms per patch, supporting real-time layer-wise inspection. Together, this work delivers a practical solution for surface quality monitoring, representing a significant advancement in intelligent quality assurance for AM processes.
    @article{Islam2026,
      title = {Beyond local receptive fields: vision transformers for real-time surface defect detection in FDM},
      volume = {143},
      ISSN = {1433-3015},
      url = {http://dx.doi.org/10.1007/s00170-026-17846-8},
      DOI = {10.1007/s00170-026-17846-8},
      number = {9-10},
      journal = {The International Journal of Advanced Manufacturing Technology},
      publisher = {Springer Science and Business Media LLC},
      author = {Islam,  MD Shafikul and Bappy,  Mahathir Mohammad and Tushar,  Saifur Rahman and Zamiela,  Christian and Wasi,  Azmine Toushik},
      year = {2026},
      month = Mar,
      pages = {5399–5420}
    }
  8. A multilayer perceptron (MLP)-based bi-directional model to predict parameters of fused deposition modeling
    Syeda Kumrun Nahar, Mohammad Muhshin Aziz Khan, Pritidipto Paul Chowdhury, Azmine Toushik Wasi, M. Morad Ali & M. Rifat Rahman
    International Journal of Advanced Manufacturing Technology (Q1) â–Ș [PDF]

    Additive manufacturing offers higher precision, improved product quality, and time-efficient operations. However, selecting optimal process parameters to achieve desired mechanical properties in the final product often relies on trial-and-error methods, contributing to material wastage. Hence, this study focused on developing a bidirectional machine learning model that integrates the final mechanical properties of PLA + parts with key fused deposition modeling (FDM) process parameters. The novelty of this work lies in the model’s bidirectional predictive capability, which allows estimation of the optimal process parameters required to achieve desired mechanical properties and, conversely, prediction of mechanical properties based on given process parameters. For this study, polylactic acid plus (PLA +) was selected. A comprehensive factorial experimental design was implemented, focusing on four critical FDM process parameters: layer thickness, printing speed, infill density, and extrusion temperature. Each parameter was examined at four levels, with each combination replicated three times, yielding 768 samples. Following ASTM standards, these samples were tested for tensile strength, flexural strength, and longitudinal shrinkage. The model exhibited high predictive accuracy, with R2 scores exceeding 0.97 and 0.99 for most parameters during the testing and training phases. Again, correlation analysis identified significant relationships between mechanical properties and input parameters. Infill density and tensile strength are positively correlated, whereas layer thickness is negatively correlated. The model’s prediction potential will enable manufacturers to make informed decisions on process parameters that will minimize material wastage and production costs and enhance sustainability.
    @article{Nahar2025,
      title = {A multilayer perceptron (MLP)-based bi-directional model to predict parameters of fused deposition modeling},
      volume = {141},
      ISSN = {1433-3015},
      url = {http://dx.doi.org/10.1007/s00170-025-16711-4},
      DOI = {10.1007/s00170-025-16711-4},
      number = {5-6},
      journal = {The International Journal of Advanced Manufacturing Technology},
      publisher = {Springer Science and Business Media LLC},
      author = {Nahar,  Syeda Kumrun and Khan,  Mohammad Muhshin Aziz and Chowdhury,  Pritidipto Paul and Wasi,  Azmine Toushik and Ali,  M. Morad and Rahman,  M. Rifat},
      year = {2025},
      month = Nov,
      pages = {3331–3346}
    }
  9. Fracture Finder: Computer-Aided Real-Time Diagnosis of Vertebral Fractures in Thoracic Spine X-Rays Using YOLOv8 with Weighted Box Fusion and Augmentation Strategies
    MD Shafikul Islam, Mahfuz Ahmed Anik, Azmine Toushik Wasi, Mahathir Mohammad Bappy
    IISE Transactions on Healthcare Systems Engineering (Q2) â–Ș [Online]

    Accurate and timely detection of vertebral-body fractures on thoracic-spine X-rays is critical, yet subtle injuries are frequently missed in high-pressure radiology workflows. We proposed Fracture-Finder, a single-stage deep-learning system that localizes suspected acute fractures with bounding boxes and returns confidence scores in real-time. Built on the YOLOv8 object-detection backbone, the model processes a typical X-ray in ≈126 ms on a mid-range consumer GPU. The model was trained and validated on 1247 expert-annotated images from the UTMB-1000 dataset. Exploratory data analysis revealed marked class imbalance and distinct trends in bounding-box size and position. To address these issues, we systematically oversampled fracture cases and applied targeted augmentations such as gamma adjustment, CLAHE, mild rotations, and grid distortion, thereby improving robustness while preserving anatomical detail. At inference, we add light test-time augmentation, along with Weighted Box Fusion, to refine box placement without sacrificing speed. On a 20 % hold-out set Fracture Finder attains mAP50=0.97, mAP50‐‐95=0.87, 93% precision and 97% recall, correctly flagging 97% of fractures with few false positives. With its high accuracy, modest hardware requirements, and rapid processing on this single-institution dataset, the model appears computationally compatible with future integration into picture archiving and communication system (PACS) workflow. Fracture-Finder offers a reliable automated tool, flagging suspicious cases for expedited review and reducing the risk of overlooked spinal injuries in both routine and emergency contexts.
    @article{Islam2026,
      title = {Fracture Finder: Computer-aided real-time diagnosis of vertebral fractures in thoracic spine X-rays using YOLOv8 with Weighted Box Fusion and augmentation strategies},
      volume = {16},
      ISSN = {2472-5587},
      url = {http://dx.doi.org/10.1080/24725579.2025.2598578},
      DOI = {10.1080/24725579.2025.2598578},
      number = {1},
      journal = {IISE Transactions on Healthcare Systems Engineering},
      publisher = {Informa UK Limited},
      author = {Islam,  MD Shafikul and Anik,  Mahfuz Ahmed and Wasi,  Azmine Toushik and Bappy,  Mahathir Mohammad},
      year = {2026},
      month = Jan,
      pages = {1–18}
    }


📔 Workshop Papers


  1. Why Benchmark Accuracy Fails to Measure Clinical Reasoning in Medical Vision Language Models: Toward Clinical Adversarial Validation
    Noor Mairukh Khan Arnob, Azmine Toushik Wasi, Mahfuz Ahmed Anik, Mohsin Mahmud Topu, Md Manjurul Ahsan
    CVPR 2026 (A*) → Medical Reasoning with VLFMs Workshop | MICAI 2026 (A) → Agentic AI for Medicine Workshop

    Recent medical vision-language models achieve strong performance on widely used benchmarks, yet it remains unclear whether these scores reflect genuine clinical reasoning or reliance on spurious correlations. This paper argues that static, correlational benchmarking is structurally insufficient for evaluating medical vision-language systems intended for high-stakes clinical use. We frame this work as a position paper and provide a principled critique of existing evaluation practices, drawing on theoretical foundations of shortcut learning and empirical evidence from recent adversarial analyses of frontier models. To address these limitations, we propose Clinical Adversarial Validation, an evaluation paradigm that compliments passive test-set assessment with clinician-guided adversarial stress testing designed to probe causal reliance, multimodal grounding, and reasoning stability. The proposed framework emphasizes counterfactual perturbations, transparent reasoning audits, and calibrated abstention as core evaluation criteria. We argue that adopting adversarial, causally informed validation is necessary to bridge the gap between benchmark performance and real-world clinical readiness.
  2. PhysLang: a Small Diagnostic Framework for Language-Grounded World Modeling
    Noor Mairukh Khan Arnob, Azmine Toushik Wasi
    ICLR 2026 (A*) → World Models Workshop

  3. Position: Biomedical NLP Demands Specialization, Not Generalization
    Azmine Toushik Wasi
    EACL 2026 (A) → Linguistic Analysis for Health Workshop â–Ș â–Ș [ACL Anthology] â–Ș

    Multimodal Artificial Intelligence (AI) promises to transform biomedicine by integrating imaging, genomics, and clinical data for superior decision-making. Yet, we contend that the current pursuit of large-scale generalist models is fundamentally misaligned with the high-risk nature of biomedical applications. This position paper argues that biomedical NLP demands specialization, not generalization, challenging the assumption that greater model scale and generality inherently ensure robustness in healthcare. We propose a theoretical framework built on three biomedical axioms: error cost asymmetry, multimodal data fragility, and interpretability–utility coupling, alongside a formal proof of criticality in biomedical NLP, showing that generalist models are intrinsically unsuited for medical tasks. As a secondary contribution, we advance a task-first design paradigm centered on modular, specialized, and ethically grounded AI architectures for biomedical use. Through analysis and illustrative cases, we contrast this approach with scale-centric strategies, exposing risks such as bias amplification, reduced interpretability, and exclusion of rare or underrepresented populations. We call for a realignment of research, funding, and regulation toward specialization as the sustainable path for meaningful and equitable biomedical AI, aiming to spark critical discourse on what constitutes genuine progress in machine learning for health.
    @inproceedings{toushik-wasi-2026-position,
        title = "Position: Biomedical {NLP} Demands Specialization, Not Generalization",
        author = "Toushik Wasi, Azmine",
        booktitle = "Proceedings of the 1st Workshop on Linguistic Analysis for Health ({H}ea{L}ing 2026)",
        month = mar,
        year = "2026",
        address = "Rabat, Morocco",
        publisher = "Association for Computational Linguistics",
        url = "https://aclanthology.org/2026.healing-1.7/",
        doi = "10.18653/v1/2026.healing-1.7",
        pages = "77--93",
        ISBN = "979-8-89176-367-8",
    }
  4. Real-World Clinical AI Requires Multimodal, Longitudinal, and Privacy-Preserving Corpora
    Azmine Toushik Wasi, Shahriyar Zaman Ridoy
    EurIPS 2025 (A) → Multimodal Representation Learning for Healthcare Workshop

  5. Differentiable Predictive Control for Precise Oxygen Level Maintenance for Critical Patients
    Azmine Toushik Wasi, Md Manjurul Ahsan
    NeurIPS 2025 (A*) → AI for Science Workshop (AI4Science)

  6. Position: Without Global Governance, AI-Enabled Biodesign Tools Risk Dangerous Proliferation
    Azmine Toushik Wasi, Mst Rafia Islam, Rahatun Nesa Priti
    NeurIPS 2025 (A*) → Biosecurity Safeguards for Generative AI Workshop

  7. Position: Adjacent Technologies Are the Key Enablers of Scalable and Safe Clinical MLLM Deployment
    Azmine Toushik Wasi, Md. Iqramul Hoque
    NeurIPS 2025 (A*) → (i) GenAI for Health, (ii) Multi-modal FMs & LLMs for Life Sciences (FM4LS), (iii) Biosecurity Safeguards for Generative AI Workshop

  8. Perspective: Lessons from Cybersecurity for Biological AI Safety and Regulation
    Azmine Toushik Wasi, Mst Rafia Islam
    NeurIPS 2025 (A*) → (i) Regulatable ML Workshop, (ii) Biosecurity Safeguards for Generative AI

  9. GFlowNets for Learning Better Drug-Drug Interaction Representations
    Azmine Toushik Wasi
    NeurIPS 2025 (A*) → Structured Probabilistic Inference & Generative Modeling

  10. Illusion of Control: Exploring the Limits of Human-in-the-Loop Oversight in Generative Finance
    Azmine Toushik Wasi, Enjamamul Haque Eram
    NeurIPS 2025 (A*) → Workshop on Generative AI in Finance

  11. Deepfakes in Political Manipulation: Evaluating Risks Under the AI Act
    Mst Rafia Islam, Azmine Toushik Wasi
    NeurIPS 2025 (A*) → Regulatable ML Workshop

  12. Seeing Isn't Believing: Addressing the Societal Impact of Deepfakes in Low-Tech Environments
    Mst Rafia Islam, Azmine Toushik Wasi
    ACM MM 2025 (A*) → Workshop on Diffusion of Harmful Content on Online Web â–Ș

  13. The Myth of ‘One Model for All’: Why Biomedical Multimodal AI Demands Specialization, Not Scale
    Azmine Toushik Wasi
    CVPR 2025 (A*) → Multimodal Foundation Models for Biomedicine

  14. A Proactive Framework for Equitable LLM Deployment in Global Health
    Azmine Toushik Wasi
    IJCAI 2025 (A*) → Multimodal Foundation Models for Biomedicine

  15. Assessing Gender Bias of Pretrained Bangla Language Models in STEM and SHAPE Fields
    Noor Mairukh Khan Arnob, Saiyara Mahmud, Azmine Toushik Wasi$
    ACL 2025 (A*) → Gender Bias in NLP

  16. Pathway-Attentive GAN for Interpretable Biomolecular Design
    Azmine Toushik Wasi, Mahfuz Ahmed Anik
    ICLR 2025 (A*) → ML for Genomics Explorations

  17. Risks and Safety Considerations for Foundation Model-based Autonomous Agents' Interaction with the Environment
    Azmine Toushik Wasi, Mahfuz Ahmed Anik, Riashat Islam
    ICLR 2025 (A*) → Foundation Models in the Wild

  18. Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems
    Mahfuz Ahmed Anik, Abdur Rahman, Azmine Toushik Wasi$#, Md Manjurul Ahsan
    NAACL 2025 (A) → LMs for Underserved Communities â–Ș

  19. Towards Culturally Inclusive Knowledge Dissemination Using AI Agents
    Mahfuz Ahmed Anik, Abdur Rahman, Azmine Toushik Wasi$#, Md Manjurul Ahsan
    AAAI 2025 (A*) → Social Impact of AI â–Ș

  20. Public Participation in AI Governance: A Meta-Analysis of Citizen Engagement Initiatives
    Mst Rafia Islam, Azmine Toushik Wasi$#
    CHI 2025 (A*) → Sociotechnical AI Governance

  21. C.A.R.E. for AI Governance: Addressing Regional Challenges in the Global South
    Mst Rafia Islam, Azmine Toushik Wasi$#
    CHI 2025 (A*) → Sociotechnical AI Governance

  22. AI for Social Good and Public Impact in Marginalized Communities: A Cross-Cultural Framework
    Azmine Toushik Wasi
    AAAI 2025 (A*) → Social Impact of AI â–Ș

  23. Deepfakes in Developing Societies: Handling the Societal Impacts and Cross-Disciplinary Vulnerabilities in Tech-Limited Environments
    Rahatun Nesa Priti, Mahir Absar Khan, Abdur Rahman, Azmine Toushik Wasi$#
    AAAI 2025 (A*) → LLM Misinformation (PDLM)

  24. Explainable Identification of Hate Speech towards Islam using Graph Neural Networks
    Azmine Toushik Wasi
    EMNLP 2024 (A*) → NLP for Positive Impact â–Ș NeurIPS 2023 → Muslims in ML â–Ș [ACL Anthology] â–Ș [arXiv] â–Ș

  25. Balancing Power and Ethics: A Framework for Addressing Human Rights Concerns in Military AI
    Mst Rafia Islam*, Azmine Toushik Wasi*$#
    AAAI 2025 (A*) → Social Impact of AI â–Ș HRAIM (Mila - Quebec) â–Ș [arXiv] â–Ș

  26. HRGraph: Leveraging LLMs for HR Data Knowledge Graphs with Information Propagation-based Job Recommendation
    Azmine Toushik Wasi
    ACL 2024 (A*) → KaLLM Workshop â–Ș [ACL Anthology] â–Ș [Youtube] â–Ș

  27. Exploring Large Language Model Systems Design Perspective Using Cognitive Ergonomics
    Azmine Toushik Wasi, Mst Rafia Islam
    EMNLP 2024 (A*) → NLP for Science â–Ș ICML 2024 → LLMs & Cognition â–Ș [ACL Anthology] â–Ș [arXiv] â–Ș [Youtube] â–Ș

  28. Exploring Bengali Religious Dialect Biases in Large Language Models with Evaluation Perspectives
    Azmine Toushik Wasi, Raima Islam, Mst Rafia Islam, TH Rafi, Dong-Kyu Chae
    CHI 2024 (A*) → HEAL Workshop â–Ș [arXiv]

  29. LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning
    Azmine Toushik Wasi, Mst Rafia Islam, Raima Islam
    CHI 2024 (A*) → In2Writing Workshop â–Ș [ACM DL] â–Ș [arXiv] â–Ș

  30. Ink and Individuality: Crafting a Personalised Narrative in the Age of LLMs
    Azmine Toushik Wasi, Raima Islam, Mst Rafia Islam
    CHI 2024 (A*) → In2Writing Workshop â–Ș [ACM DL] â–Ș [arXiv] â–Ș

  31. SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks
    Azmine Toushik Wasi, MD Shafikul Islam Sohan, Adipto Raihan Akib
    AAAI 2024 (A*) → GCLR Worksop â–Ș [Paper Site] â–Ș [arXiv] â–Ș [GitHub] (~100⭐)

  32. Optimizing Inventory Routing: A Decision-Focused Learning Approach using Neural Networks
    MD Shafikul Islam Sohan, Azmine Toushik Wasi
    NeurIPS 2023 (A*) → New in ML â–Ș [Paper Site] â–Ș [arXiv] â–Ș


📑 Challenges & Shared Tasks


  1. CIOL at AraGenEval shared task: Authorship Identification and AI Generated Text Detection in Arabic using Pretrained Models
    Sadia Tasnim Meem, Azmine Toushik Wasi
    EMNLP 2025 (A*) → ArabicNLP 2025 â–Ș

  2. CIOL at SemEval-2025 Task 11: Multilingual Pre-trained Model Fusion for Text-based Emotion Recognition
    Md. Iqramul Hoque, Mahfuz Ahmed Anik, Abdur Rahman, Azmine Toushik Wasi$
    ACL 2025 (A*) → SemEval-2025 â–Ș

  3. CIOL at CLPsych 2025: Using Large Lanuage Models for Understanding and Summarizing Clinical Texts
    Md. Iqramul Hoque, Mahfuz Ahmed Anik, Azmine Toushik Wasi$
    NAACL 2025 (A) → CLPsych â–Ș

    The increasing prevalence of mental health discourse on social media has created a need for automated tools to assess psychological wellbeing. In this study, we propose a structured framework for evidence extraction, well-being scoring, and summary generation, developed as part of the CLPsych 2025 shared task. Our approach integrates feature-based classification with context-aware language modeling to identify self-state indicators, predict well-being scores, and generate clinically relevant summaries. Our system achieved a recall of 0.56 for evidence extraction, an MSE of 3.89 in well-being scoring, and high consistency scores (0.612 post-level, 0.801 timeline-level) in summary generation, ensuring strong alignment with extracted evidence. With an overall good rank, our framework demonstrates robustness in social media-based mental health monitoring. By providing interpretable assessments of psychological states, our work contributes to early detection and intervention strategies, assisting researchers and mental health professionals in understanding online well-being trends and enhancing digital mental health support systems.
    @inproceedings{hoque-etal-2025-ciol,
        title = "{CIOL} at {CLP}sych 2025: Using Large Lanuage Models for Understanding and Summarizing Clinical Texts",
        author = "Hoque, Md. Iqramul  and Anik, Mahfuz Ahmed  and Wasi, Azmine Toushik",
        booktitle = "Proceedings of the 10th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2025)",
        month = may,
        year = "2025",
        address = "Albuquerque, New Mexico",
        publisher = "Association for Computational Linguistics",
        url = "https://aclanthology.org/2025.clpsych-1.19/",
        doi = "10.18653/v1/2025.clpsych-1.19",
        pages = "235--241",
        ISBN = "979-8-89176-226-8",
      }
  4. HerWILL@DravidianLangTech 2025: Ensemble Approach for Misogyny Detection in Memes Using Pre-trained Text and Vision Transformers
    Neelima Monjusha Preeti, Trina Chakraborty, Noor Mairukh Khan Arnob, Saiyara Mahmud, Azmine Toushik Wasi$
    NAACL 2025 (A) → DravidianLangTech â–Ș

  5. NLPopsCIOL@DravidianLangTech 2025: Classification of Abusive Tamil and Malayalam Text Targeting Women Using Pre-trained Models
    Abdullah Al Nahian, Mst Rafia Islam, Azmine Toushik Wasi$#, Md Manjurul Ahsan
    NAACL 2025 (A) → DravidianLangTech â–Ș

  6. Eureka-CIOL@DravidianLangTech 2025: Using Customized BERTs for Sentiment Analysis of Tamil Political Comments
    Enjamamul Haque Eram, Anisha Ahmed, Sabrina Afroz Mitu, Azmine Toushik Wasi$#
    NAACL 2025 (A) → DravidianLangTech â–Ș

  7. MysticCIOL@DravidianLangTech 2025: A Hybrid Framework for Sentiment Analysis in Tamil and Tulu Using Fine-Tuned SBERT Embeddings and Custom MLP Architectures
    Minhaz Chowdhury, Arnab Laskar, Taj Ahmad, Azmine Toushik Wasi$#
    NAACL 2025 (A) → DravidianLangTech â–Ș

  8. Akatsuki-CIOL@DravidianLangTech 2025: Ensemble-Based Approach Using Pre-Trained Models for Fake News Detection in Dravidian Languages
    Mahfuz Ahmed Anik, Md. Iqramul Hoque, Wahid Faisal, Azmine Toushik Wasi$#, Md Manjurul Ahsan
    NAACL 2025 (A) → DravidianLangTech â–Ș

  9. IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages
    Siddhant Gupta*, Siddh Singhal*, Azmine Toushik Wasi*$#
    COLING 2025 (B) → NLU of Devanagari Script Languages (CHiPSAL) â–Ș

  10. RoBERTa Ensemble for Identifying Children’s Medical Disorders in English Tweets
    Azmine Toushik Wasi, Sheikh Ayatur Rahman
    ACL 2024 (A*) → SMM4H Workshop Task: 5 â–Ș [ACL Anthology] â–Ș [Youtube] â–Ș

  11. Analyzing Social Anxiety Effects through Context-Aware Transfer Learning on Reddit Data
    Sheikh Ayatur Rahman, Azmine Toushik Wasi
    ACL 2024 (A*) → SMM4H Workshop Task: 3 â–Ș [ACL Anthology] â–Ș [Youtube] â–Ș â–Ș

  12. ML Algorithm Synthesizing Domain Knowledge for Fungal Spores Concentration Prediction
    Md Asif Bin Syed, Azmine Toushik Wasi, Imtiaz Ahmed
    IISE 2023 → QCRE Data Challenge 2023 â–Ș [Site] â–Ș [Technical Report] â–Ș [Code] â–Ș â–Ș


📃 Pre-prints / Works in Review

  1. Retrieval Geometry Shapes Cache-Based Clip Adaptation
    Mahir Shahriar Tamim*, Md. Samiul Alim*, Azmine Toushik Wasi, Shahriyar Zaman Ridoy, Meharun Nesa, M. Abu Yousuf, Alex Lamb, M. Ali Moni
    [arXiv]

    Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image–image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, ImageNet-A cache gain ranges from at most +0.44 points for CLIP and MAE to +19.7 ± 0.4 for DINOv2-L, while label-free retrieval-space selection retains 98% of oracle gain on ImageNet-V2. These results show that memory quality depends not only on which examples are stored, but also on how they are retrieved. Motivated by this finding, we propose MARC (Memory Augmented Retrieval for CLIP), a training-free system that uses frozen CLIP for prediction and DINOv2-B for retrieval with a single fusion weight. A single-view cache repairs 1,074 ± 21 baseline errors, compared with 878 ± 4 for a 64-view ensemble, at roughly one seventh of the cost. Across four ImageNet distribution shifts, MARC reaches a 67.91% OOD average and, at matched DINOv2-B scale and eight views, achieves 64.17 ± 0.31% versus 62.75 ± 0.15% for a graph-based cache system while running 2.6 times faster. Overall, our results establish retrieval space as a first-order design choice for robust cache-based adaptation in remote sensing, scientific imaging, and changing visual environments.
    @misc{tamim2026retrievalgeometryshapescachebased,
          title={Retrieval Geometry Shapes Cache-Based Clip Adaptation}, 
          author={Mahir Shahriar Tamim and Md. Samiul Alim and Azmine Toushik Wasi and Shahriyar Zaman Ridoy and Meharun Nesa and Mohammad Abu Yousuf and Alex Lamb and Mohammad Ali Moni},
          year={2026},
          eprint={2609.23409},
          archivePrefix={arXiv},
          primaryClass={cs.CV},
          url={https://arxiv.org/abs/2609.23409}, 
    }
  2. InvariantBench: Can Large Language Models Exhibit Inherent Reasoning Consistently Across Equivalent Transformations?
    Azmine Toushik Wasi, Mahir Absar Khan, Abdur Rahman, Wahid Faisal, Sukanta Saha, Saimon Bhuiyan, Mahdiya Rahman Sukanya, Rahatun Nesa Priti, Md. Iqramul Hoque, Munem Shahriar, Raima Islam, Sanatan Sushil, Shahriyar Zaman Ridoy, Kazi Rajwan Sultan, Md Tanzib Hosain, MD Shafikul Islam, Dong-Kyu Chae, Md Manjurul Ahsan, Md Rizwan Parvez

    Reasoning is often attributed to large language models (LLMs), yet it remains unclear whether they operate over underlying semantics or rely on surface-form patterns. Existing benchmarks evaluate correctness on fixed problem instances, but overlook a fundamental property of reasoning: invariance under semantics-preserving transformations. If a model truly understands a problem, its predictions should remain consistent across equivalent representations. We introduce InvariantBench, a benchmark of 1159 seed problems invariant forms, spanning 16 tasks across 3 reasoning families and 12 fine-grained invariance axes. Each problem is paired with multiple semantically equivalent variants under a strict invariance contract, enabling evaluation beyond accuracy to measure consistency across representations. Experiments on 15 frontier and open-weight LLMs reveal a persistent invariance gap: invariant reformulations reduce accuracy by up to 30% for strong models, while the gap between solving at least one versus all four variants reaches 60%. Full consistency remains below 5% for most open-weight models, compared with ~48% for the two strongest models and 98.6% for human experts. These results show that high base accuracy substantially overestimates reasoning ability, and establish invariance as a necessary axis for evaluating and improving robust language understanding.
  3. Characterizing Errors in Small Medical Language Models: Geometry, Prediction, and Causal Sensitivity
    Azmine Toushik Wasi, Adinath Madhavrao Dukre, Md Manjurul Ahsan, Muhammad Awais, Sara Atito

    What do differences between correct and incorrect medical answers reveal about a language model's errors? Geometric separation, predictive internal signals, and intervention effects offer different evidence, yet none alone establishes that a model recognizes or can correct its mistakes. We introduce Mirage, a framework that tests these claims separately on shared questions and responses, using coverage checks, predictive baselines, and controlled interventions. Across five medical checkpoints and four datasets, we collect 208,896 responses across calibration and evaluation. Incorrect responses are more dispersed in 545 of 808 eligible evaluation pairs, but this geometric comparison covers only 12.6% of the 6,400 evaluated pairs. On 278 matched pairs, masking explicit answers reduces biomedical Fisher label propagation macro-F1 from 0.807 to 0.670, revealing a substantial contribution from answer content. Internal states predict errors on unseen questions, while gains beyond answer type vary across visual tasks and fitting partitions. Interventions produce protocol dependent effects, and a supplementary restoration study finds no observed benefit in sampled answer recovery. Together, these results identify three obstacles to interpreting correctness signals: selective coverage, dependence on answer content and task structure, and a gap between output sensitivity and demonstrated repair. By turning these obstacles into explicit testbeds, we provide a reproducible framework for testing what geometric, predictive, and intervention evidence can establish about model errors, while keeping benchmark agreement separate from clinical explanation.
  4. Why Do Finetuned ECG Foundation Models Fail at New Hospitals? Separating What Finetuning Finds from What It Adds
    Srikar Vardhan, Adinath Madhavrao Dukre, Azmine Toushik Wasi, Imran Razzak

    Finetuned ECG foundation models often perform worse at hospitals other than the one where they were trained, but which part of the finetuned model fails is rarely examined. We decompose the gain from finetuning, for each diagnosis, into \emph{finding}, information that the frozen pretrained model already holds but that labels alone do not locate, and \emph{adding}, information that the frozen model does not contain. Finding is measured by distilling the finetuned model into a readout of the frozen layers, and the decomposition passes three preregistered validity tests. We study three open ECG foundation models (HuBERT-ECG, ECG-FM, ST-MEM), each finetuned with three seeds on 19 diagnoses from PTB-XL, and evaluate them on five external cohorts. Most of the gain is finding. Adding is significant in all 14 cases of a core of infarct, conduction and normal diagnoses selected on validation data, but in only 3 of 16 other cases. At new hospitals, adding retains a median 37% of its value against 69% for finding, while most failures come from decision thresholds, with specificity as low as 0.40, and from differing label definitions. Three remedies hold under preregistered tests. Averaging the finetuned output with a linear readout of the frozen layers raises external sensitivity at 90% specificity from 76.5% to 78.5% without site labels or retraining. Freezing the first four transformer blocks improves hypertrophy detection in two models. Recalibrating thresholds on 100 verified local negatives restores the intended specificity in all 111 cases tested.
  5. Think Before You Locate: Clinical Chain-of-Thought Supervision for Sequential Medical Image Grounding
    Adinath Madhavrao Dukre, Govinda Kolli, Yifan Lu, Ziyun Zou, Azmine Toushik Wasi, Sara Atito, Behzad Bozorgtabar, Dwarikanath Mahapatra, Muhammad Awais, Imran Razzak

    Clinicians rarely localize a finding from one image alone. They compare current and prior scans, follow structures across slices, and match regions across views or modalities. Yet vision-language models for sequential medical grounding usually learn only from bounding boxes. Boxes show where the target is, but they do not explain why that region in that image answers the question. We study whether explicit, image-grounded reasoning can improve localization, and which tasks benefit from it. We introduce \method, whose core is MedChain, a corpus of 182,573 filtered rationales spanning eight sequential grounding tasks and ten imaging modalities. A medical VLM (FlemingVL-38B) first drafts a caption-like rationale, and a multimodal refiner (Qwen3.5-397B-A17B) then re-examines the images, question and annotated box to write a box-conditioned explanation with explicit cross-image reasoning. We fine-tune Qwen3.5 on these rationales with deliberative chain-of-thought supervision (dCoT-SFT). Across 4B, 9B, and 27B models, dCoT-SFT sets a new state of the art on MedSG-Bench. The 27B model reaches 77.24 mIoU and 85.20 Acc@0.5, improving over the strongest prior model by 4.69 and 5.49 points. In a paired comparison using fixed 4B checkpoints, dCoT-SFT improves over answer-only SFT by 3.56 mIoU (95\% CI 3.13--4.00). The gains mainly come from change detection and object tracking, while performance drops on three appearance-matching tasks. Replacing a model's own rationale with a rationale from another task reduces mIoU by 12.2 points, showing that localization depends strongly on the reasoning context. Finally, adding GRPO with an IoU reward after dCoT-SFT gives only small and task-dependent gains, which we analyze through controlled multi-seed development experiments.
  6. FrontierSpatial: Holistic Evaluation of Spatial Reasoning in Large Multimodal Models
    Wahid Faisal*, Azmine Toushik Wasi*, Mohammed Eunus Ali, Md Rizwan Parvez

    Spatial reasoning is foundational to intelligent systems, yet its evaluation in multimodal large language models (MLLMs) remains fragmented across narrow benchmarks that test isolated skills. We introduce FrontierSpatial, a unified benchmark and evaluation framework for large-scale spatial reasoning. We consolidate 23 existing datasets into 7,838 questions paired with 11,825 images, organized into 7 categories, 19 sub-categories, and 65 fine-grained tasks spanning physical, geometric, knowledge-grounded, relational/causal, temporal, quantitative, and perceptual reasoning. Beyond the benchmark, we provide (1) an automated curation pipeline that distills over 60,000 raw instances into a balanced dataset and generalizes to other domains, (2) an automated evaluation harness for reproducible assessment of 16 state-of-the-art MLLMs, and (3) an automated error-analysis harness for fine-grained diagnosis across the task hierarchy. Our evaluation reveals large and uneven performance gaps across spatial skills, exposing systematic weaknesses that remain hidden in isolated benchmarks. FrontierSpatial provides a rigorous foundation for systematic, fine-grained evaluation and deeper understanding of multimodal spatial reasoning.
  7. How Far Can Instruction Steering Go? Coverage, Quality, and Calibration
    Shubhashis Roy Dipta, Azmine Toushik Wasi, Sukanta Saha, Sanatan Sushil, Md. Masudur Rahman

    Activation steering promises modular control of language models: learn one direction per instruction, then add the directions that a request needs. This promise rests on two assumptions: directions that work alone also work together, and an answer that passes an instruction checker is still a good answer. We test both assumptions on seven models. We calibrate one edit per instruction and freeze it. An answer counts as a success only if it passes every checker and a blind quality judge, which agrees with a human audit. We separate installation, where the edit must carry an instruction that the prompt omits, from reinforcement, where the prompt already states it. Few instructions can be installed reliably. The quality judge shrinks this set further. On four smaller models, it removes all four installations that the automatic checks accept, and it roughly halves the usable reinforcements. Where installations do succeed alone, as on Gemma-3-27B, they also compose. On instruction pairs, they beat the unedited model by 35.0 points. Still, they trail simply stating the instructions by 17.2 points. Reinforcement gives mixed results and falls behind prompting as sets grow, yet much of this loss comes from calibration. Composing steering directions is therefore feasible, but its value depends on coverage, answer quality, and calibration, and prompting remains the baseline to beat.
  8. DeepVisualMath: Comprehensive Evaluation of Visual Mathematical Reasoning in Large Multimodal Models
    Azmine Toushik Wasi*, Wahid Faisal, Mohammed Eunus Ali*, Md Rizwan Parvez

    Visual mathematical reasoning requires integrating visual evidence with mathematical concepts, yet its evaluation in multimodal large language models (MLLMs) remains fragmented across benchmarks targeting limited domains and skills. We introduce DeepVisualMath, a unified benchmark and evaluation framework for holistic assessment of visual mathematical reasoning. The benchmark comprises 10,298 questions, each paired with one image, organized into 7 categories, 18 subcategories, and 80 tasks, spanning applied sciences, continuous mathematics, relational geometry across dimensions, discrete mathematics and formal logic, plane geometry, contextual mathematics, and solid geometry. Beyond the benchmark, we provide an automated curation pipeline, a reproducible evaluation harness, and an automated error analysis framework for detailed diagnosis across the task hierarchy. Evaluation of 10 MLLMs reveals substantial and uneven performance gaps: overall accuracy ranges from 15.50% to 66.67%, with even the strongest model leaving approximately one third of questions unresolved. Performance also varies markedly across mathematical domains; the leading model achieves 77.36% in continuous mathematics but only 51.85% in plane geometry, highlighting limitations obscured by aggregate scores. These findings demonstrate that strong performance in one mathematical domain does not imply broad visual mathematical competence. DeepVisualMath provides a systematic foundation for evaluating these capabilities and identifying priorities for improving mathematical reasoning grounded in visual evidence.
  9. BRACE: An R Benchmark for Rigorous Biomedical Data Science Code Search
    Md Tanzib Hosain, Azmine Toushik Wasi, Salman Rahman, Md Mofijul Islam, Mohammad Ali Moni, Mohammed Eunus Ali, Usman Naseem, Md Rizwan Parvez

    Efficient code retrieval is critical for biomedical data scientists, who navigate thousands of R packages on CRAN and Bioconductor and a growing body of open-source analysis pipelines. Existing code search benchmarks, however, focus on Python and rarely stress-test robustness beyond superficial lexical cues. To address this gap, we adapt an automated benchmark-construction pipeline to R and present BRACE (Biomedical R Anonymized Code rEtrieval), a benchmark built from real-world biomedical repositories and packages. BRACE contains 1,208 query-code pairs for evaluation and 5,310 pairs for training. Queries are LLM-generated natural-language descriptions validated through scoring by domain experts and hypothesis testing. The pipeline first ensures that every snippet byte-compiles and resolves all of its dependencies inside a pinned R/Bioconductor environment. It then categorizes snippets by dependency complexity, distinguishing functions that use only base R, functions that rely on custom S3/S4/R6 or Bioconductor container classes, and functions that invoke user-defined helpers. BRACE further stress-tests retrieval robustness through identifier anonymization and through representation shift, in which functions are lowered to R bytecode or serialized as S-expression parse trees. Under these conditions, our evaluation of six retrieval models reveals consistent drops when identifiers are anonymized, NDCG falls by 3-25 points on base-R functions and by up to 56 points on container-class functions, and far larger drops on lowered representations, where even the strongest model ranks the correct function first for fewer than a quarter of bytecode queries. The results indicate that current models still rely on lexical features rather than code semantics, and that in R this reliance extends to string-typed data accessors.
  10. Characterizing Age Conditioned Affective Reasoning in Vision Language Models
    Shahriyar Zaman Ridoy, Azmine Toushik Wasi, Lin Gu, Ruogu Fang

    Vision language models now support personalized and emotionally aware interactions, but demographic conditioning may reflect stereotypes rather than real human variation. We test this for age-conditioned affect using all 900 OASIS images and 17 proprietary and open-weight VLMs under no-age, young, middle, and older-adult conditions. We collect valence and arousal ratings, emotion labels, and open-ended descriptions, and compare them with participant-level judgments from 822 human raters. Models show high baseline competence (median valence $r = 0.891$) and strong age responsiveness: 16 of 17 lower valence and all 17 lower arousal for older personas. Yet image-specific fidelity remains near zero ($r_{\text{part}} = -0.030$ to $0.093$, ceiling $0.86$), even though model age effects agree with one another (mean pairwise $r = 0.428$). The mismatch is content-dependent: on pleasant sensitive images, models lower older-persona valence by $0.845$ points while older human raters shift in the opposite direction ($+0.443$), and older-persona descriptions use more memory and vulnerability language. We also test six agent-based interventions on 300 images; reasoning alone does not improve alignment, while human grounding improves the broad direction of age effects but leaves substantial image-specific error, with the best agent reaching $r_{\text{part}} = 0.27$. These results show that matching average ratings or producing a plausible age persona is not enough: faithful personalization requires models to change on the same stimuli, and in the same direction, as the human group they aim to represent.
  11. VISTA: Combating Visual Token Drift in Long-Form Video Language Generation via Saliency-Gated Attention Refresh
    Azmine Toushik Wasi, Ankan Deria, Adinath Madhavrao Dukre, Sara Atito, Muhammad Awais, Imran Razzak

    Long-form video-language generation requires sustained visual grounding over hundreds of decoding steps, yet decoder-only video-language models progressively lose access to the input video as generation continues. We identify this failure mode as \emph{visual token drift}: as the autoregressive KV cache grows, attention to the fixed visual-token prefix is diluted, causing later tokens to rely more on language priors than visual evidence. We present VISTA (\textbf{V}isual \textbf{I}nference with \textbf{S}aliency-gated \textbf{T}oken \textbf{A}ttention), a training-free inference-time framework that mitigates this effect through three complementary mechanisms: periodic visual context refresh to re-anchor decoding to the original multimodal input, attention-rollout saliency scoring to identify under-attended frames without gradients, and GQA-aware key boosting to amplify neglected visual tokens in grouped-query attention decoders. VISTA requires no fine-tuning, no architectural changes, and introduces modest inference overhead. On VideoChatGPT across the Generic and Temporal subsets, VISTA improves ROUGE-L, CIDEr, BERTScore, and lexical diversity over standard greedy decoding, with the largest gains on temporally demanding examples. These results establish visual token drift as a key inference-time bottleneck in long-form video-language generation and show that it can be mitigated through lightweight decoding-time interventions.
  12. Memory and Wealth Debt in Sequential Monitoring
    Azmine Toushik Wasi, Shubhashis Roy Dipta

    Anytime-valid monitors control false alarms throughout deployment, yet they can respond slowly to changes after a long compliant history. We explain this delay through two statistical costs: estimator memory, which keeps stakes tied to old observations, and wealth debt, the accumulated loss of log evidence that later bets must recover before an alarm. We prove a finite-horizon lower bound on the log wealth of windowed exponential betting that holds for arbitrary window contents and wealth at the change, separating memory, estimation error, debt, and noise. From this bound, we derive conditions for e-BH to recover all affected subgroups and show how mixing over start times avoids earlier debt at an explicit evidence cost while preserving lifetime validity. Simulations with 10,000 compliant batches reveal the resulting trade-off: short windows accumulate debt, while long windows adapt slowly. At the null boundary, the start mixture detects 96.3\% to 97.9\% of changes across all tested launch offsets, compared with about 62\% for windowed betting alone. On resampled real ACS-Income losses with an injected risk increase of 0.05, cumulative exponential betting detects none of 200 shifts, while windowed betting and the start mixture detect all 200. These results provide a statistical basis for designing monitors that retain sensitivity to emerging risks after prolonged deployment while maintaining false alarm control, supporting timely intervention when system performance deteriorates.
  13. When Do Subgroup Specific Acceptance Policies Help in Medical Evidence Verification?
    Azmine Toushik Wasi, Md Tanzib Hosain, Adinath Madhavrao Dukre, Md Manjurul Ahsan, Muhammad Awais, Sara Atito, Imran Razzak

    Medical evidence verification requires acceptance decisions under unequal error costs and uncertain score reliability. Category-specific rules offer flexibility but can increase estimation error when calibration data are limited. We formalize this trade-off by decomposing risk into population gains and estimation error, separating cost heterogeneity from score reliability. Across six language models and four benchmark collections, we compare shared, partially pooled, and category-specific policies using source-grouped evaluation, refit bootstrapping, and a 216-cell simulation. Category-specific empirical risk minimization lowers point-estimate cost in 17 of 24 settings, yet all corresponding refit intervals include zero. Shared calibration with category-dependent costs performs better in 8 of 12 biomedical settings. Our audits reveal sensitivity to small samples, category imbalance, and scoring conventions. These findings support shared calibration with explicit error costs as a starting point for medical answer acceptance and review, with category-specific rules requiring held-out gains beyond fitting uncertainty.
  14. What ECG Representations Discard: Amplitude Loss, Recoverability, and Transfer Across Hospitals
    Srikar Vardhan Mangadoddi, Adinath Madhavrao Dukre, Azmine Toushik Wasi, Imran Razzak

    Model failures across hospitals are often attributed to input distribution shift. We instead study what representations discard, using label-free linear probes to audit whether frozen encoders retain the physiological quantities that define diagnoses. Standard per-lead ECG normalization removes amplitude needed for voltage diagnoses: voltage recoverability falls from $R^2=0.97$ on raw waveforms to $0.16$ after normalization. Restoring amplitude improves transfer; a random-feature control does not. Recoverability separates diagnoses that transfer from those needing repair. For the PR interval, three encoders transfer in the order of their recoverability, linking the two within a task. We also show how local labeling thresholds degrade operating points while preserving rankings, a concept shift that recalibration repairs. Four hospitals agree on timing thresholds but differ on voltage criteria, establishing this failure mode without making it dominant in our data. We combine these signals in TRACE, a cost-aware audit that routes diagnoses to shipping, recalibration, detector substitution, feature restoration, probing, or refusal. Across three held-out hospitals and two encoders, TRACE achieves at least 95\% of oracle utility in all sixteen routing cells; no fixed policy does. We preregistered 38 predictions with numerical pass and kill conditions before measurement.
  15. PAT: Parallel Agent Tuning for Coordinator-Free Plug-and-Play Multi-LLM Training with Monotonic Improvement Guarantees
    Md Tanzib Hosain, Azmine Toushik Wasi, Anirban Saha Anik, Md Rizwan Parvez

    Teams of small language models can match or outperform a single large model, but jointly updating multiple agents introduces compounding distribution shifts and inter-agent coupling, and existing remedies serialize training so that the cost of a stage grows with the number of agents. We introduce Parallel Agent Tuning (PAT), a coordinator-free, synchronous training paradigm that represents the team as a factorized policy and updates all agents simultaneously from a single shared on-policy batch, without a central controller or an update schedule. PAT couples a shared on-policy advantage estimator with per-agent, per-state KL trust regions whose additivity yields a team-level budget, and it characterizes the interaction gap between the team's joint surrogate and the agents' independent surrogates, which is either budgeted explicitly (PAT-I) or eliminated through a joint sequence ratio (PAT-J). The framework guarantees monotonic improvement at every synchronous stage and plug-and-play invariance: any agent can be upgraded to a stronger model without retraining the rest of the team while the performance bound improves. Under a matched wall-clock budget on identical hardware, a team of three 4B agents trained with PAT surpasses Qwen3-32B on AIME24/25 by 5.6 points on average and the same team trained by sequential agent tuning by 1.8 composite points, because PAT completes 1.6$\times$ as many certified stages in the same time; at matched stages the two agree within seed variance. Swapping in two 8B agents boosts the composite by 5.6.
  16. Parallel Knowledge Acquisition: MAS Pedagogical Strategies for LLMs' Concurrent Ontology Learning
    Md Tanzib Hosain, Azmine Toushik Wasi, Anirban Saha Anik, Md Rizwan Parvez

    Dialogic pedagogy has shown that a learner large language model (LLM) can acquire a verifiable ontology from a single teacher LLM, but that paradigm is dyadic and serial: one omniscient teacher, one learner, one channel. Human learning is rarely organised this way; knowledge is distributed across several more-knowledgeable others, acquired concurrently, and consolidated with peers. We introduce \emph{Parallel Knowledge Acquisition} (PKA), in which a learner acquires a target ontology through several concurrent pedagogical channels, teachers holding partial views, learners exchanging intermediate summaries, or both, and must integrate the streams into one representation. We formalise a PKA configuration as a topology, a knowledge partition, a schedule, a budget, and a per-channel strategy, and evaluate 48 configurations spanning fan-in, fan-out with peer channels, and jigsaw meshes against dyadic and serial-partial controls at matched budget, on synthetic ontologies of alien species (fifty entity--feature--value triplets) with GPT-4o and Llama-3.3-70B. Learners are assessed by reconstructing the ontology and by playing the 20 Questions Game. Two-teacher fan-in with top-down channels retains 97\% of dyadic accuracy; concurrent and serial delivery of the same partial views differ by 0.3 triplets when the learner can consult transcripts and by 6.2 when it must integrate from its own notes. By-entity partitions raise bottom-up instruction from 21.5 to 38.2 of 50 triplets when per-channel budget is preserved. Peer summaries cut the variance of learner-led strategies by 53\%; second-hand knowledge in jigsaw meshes is recovered at 42\% of first-hand fidelity, 68\% with peer questions. The dissociation between reconstruction and 20 Questions efficiency persists in every topology ($r=-0.07$).
  17. Last Translation Benchmark
    VilĂ©m Zouhar, Niyati Bafna, Mukund Choudhary, Maike ZĂŒfle, ... , Azmine Toushik Wasi, ... (and 255 others)

    For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulnerable to reward-hacking. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.
    @misc{zouhar2026translationbenchmark,
          title={Last Translation Benchmark}, 
          author={VilĂ©m Zouhar and Niyati Bafna and Mukund Choudhary and Maike ZĂŒfle and Sara Rajaee and Pinzhen Chen and Jannis Vamvas and Sara Papi and Ona de Gibert and Bhavitvya Malik and Eliya Habba and Orfeas Menis Mastromichalakis and PatrĂ­cia SchmidtovĂĄ and Michelle Wastl and Sheriff Issaka and Leshem Choshen and Stella Biderman and Antonis Anastasopoulos and Jan Niehues and Rico Sennrich and Mrinmaya Sachan and Ondƙej Bojar and Kenton Murray and Jörg Tiedemann and Alham Fikri Aji and Philipp Koehn and Christof Monz and Alexandra Birch and Sowmya Vajjala and Chalamalasetti Kranti and Cristina España-Bonet and Nobin Sarwar and David KaczĂ©r and Sourajit Saha and Jonathan Tonglet and Shunta Asano and Malik Marmonier and Daban Q. Jaff and Vaisakhi Mishra and Hend Al-Khalifa and Gabriele Sarti and Nils Rehlinger and Juan Daniel Cuervo Villa and Dominik Macháček and Saugata Purkayastha and Jagannathan Ramanujam and Shubhashis Roy Dipta and Aviral Nigam and Shuaib Shuaib Yusuf and Heejin Do and Jonathan Yahav and Johannes-Rudolf David and Maria Carmen Staiano and Zuzana Nadova and Fred Philippy and Ron Keinan and Maria Lymperaiou and Silvia Casola and Fabian Retkowski and AndrĂ©s Jerez and Hanna Yukhy- menko and Sangwon Ryu and Avantica Vempati and Sukannya Purkayastha and Adrian Cosma and Erivan Inan and Vitalii Babenko and Wafa Aissa and Valentin Scourneau and Fatima Haouari and Venkata Prasanth Kumar Gummadi and Mehdi Jafarzadeh and Manon Reusens and Kaiser Sun and Lukas Edman and Shaomu Tan and Giuseppe Gallipoli and Pawan Sasanka Ammanamanchi and Manar Ali and Mohammad Sadegh Gholizadeh and Dipankar Srirag and Marek Ć uppa and Javier GarcĂ­a Gilabert and Ruta Binkyte and Ana-Maria Bucur and Sabry E. Farrag and Youssef Saber and Yihong Liu and Theresia Veronika Rampisela and Christian Hoang and Nicoleta Cojocaru and Jan KocoƄ and Jean Maillard and Xiaochuang Yuan and Sina Ahmadi and Daryna Dementieva and Philipp Mondorf and Kaustubh Dhole and Valmik Nahata and Roman Wixinger and Amir Hossein Yari and Shenbin Qian and Manuel Tuor and Fida Mohammad Thoker and Sergey Troshin and Lance Calvin Lim Gamboa and Amir Arsalan Rezapour and KĂ€triin Kukk and Koel Dutta Chowdhury and Shaswati Saha and Seth Aycock and Bo Chen and Linh Vu and Vatsal Venkatkrishna and Shayan Bali and Arafat Ahsan and Luan Thanh Nguyen and Hassan Soliman and Ngoc Quynh Tram Do and Azmine Toushik Wasi and R. Damanhuri and Marius Huber and Kazuki Egashira and Jimson Paulo Layacan and David Africa and Vladislav Poritski and Mike Zhang and Deep Shah and Abdulaziz Nura Kani and Luis Frentzen Salim and Paul Gavrikov and Bello Umar Bello and Ayush Sunil Munot and Anumit Garg and Yolanda Xavier and Qiaoyuan Zheng and Kawsar Ahmed and Debanshu Das and Zimu Wang and Gengyu Rao and Kamile Dementaviciute and Farhan Farsi and L D M S Sai Teja and Dawei Zhu and Yi Fan and Wei Liu and Marco Gaido and Elias Herranen and Sankalan Pal Chowdhury and Karen Sanchez and Guy Kaplan and Farzad Shami and Ashok Urlana and Amir Hossein Kargaran and Sofie Goethals and Priyaranjan Pattnayak and Oksana Volchek and Marii Ojastu and Hongbin Na and Emilian Radoi and Chenyi Zhao and Carlos Hinojosa and Andrei Niculae and Andrea Gregor de Varda and Zaid Alyafeai and Tomasz Limisiewicz and Reem Alzahrani and Pouya Sadeghi and Nehal Kathrotia and Mateusz Lango and Enzo Doyen and Alex FlĂŒckiger and Ulysses Sekai Tully Carr and Samuel Simko and Ritwik Tiwari and Rishit Dagli and Isaac R Caswell and Bowen Yi and Aicha Chorana and Selja KerĂ€nen and Sadiksha Chitrakar and Muhammad Ravi Shulthan Habibi and Joy Olusanya and Bishal Shrestha and Zhengxiang Wang and Vivek Harsha Lakkamaneni and Sophia Conrad and Panayiotis Panayiotou and Nazia Tasnim and Marta Punsola MunĂĄrriz and Marko Culjak and Luis Lara and Jenny Chim and Jannatul Nayem and Fidel RodrĂ­guez VelĂĄsquez and Eran Yahav and Blanka KövĂ©r and Beatrice Savoldi and Anmol Goel and Aishik Mandal and Tosin Adewumi and Raoyuan Zhao and Mykola Haltiuk and Antonia Karamolegkou and Yuxing Lu and Thura Aung and Naser Almousa and Tommaso Cerruti and Raia Abu Ahmad and Beni Egressy and Alireza Pakniat and StĂ©phane J. P. S. Thunus and Rachel Bawden and Lena Libon and Samridhya Biswas and Prakhar Gupta and Nusrat Jahan Lia and Nguyen Tai and Natchapon Jongwiriyanurak and Minh Ngoc Do and Ivan Barać and Dzmitry Kuzmin and Badal Nyalang and Antoine Taroni and Andy Catruna and Rushikesh Zawar and Roland Aydin and Pavel Stepachev and Ilai Yaron Levy and Andreas Simons and Rayyan Merchant and Ziyi Yang and Samuel Frontull and Kenneth Enevoldsen and Harris Abdul Majid and Tim Graf and Tatiana Bielakova and Sharifa Djurabaeva and Shaoxiong Ji and Jirui Qi and Ayla Rigouts Terryn and Yurii Paniv and Xiyan Fu and Sunisth Kumar and Shree Harsha Bokkahalli Satish and Papa Abdou Karim Karou Diallo and Mengyu Ye and Maximilian Vieweg and Matija Akrap and KristĂœna OnderkovĂĄ and Joseph Attieh and Ivan Yuri De Leon and Ibrahim Baroud and Esrael Teferi Tensay and Elisabeth Fittschen and David Dukić and BenoĂźt Sagot and Jingwei Ni and Yu Fan and Juri Opitz},
          year={2026},
          eprint={2609.04173},
          archivePrefix={arXiv},
          primaryClass={cs.CL},
          url={https://arxiv.org/abs/2609.04173}, 
    }
  18. Safe Decision Making with Distribution-Free Per-Group Utility Guarantees by Risk-Averse Calibration
    Azmine Toushik Wasi, Md Manjurul Ahsan

    Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard guarantees are typically marginal and can mask systematic undercoverage within demographic or other subgroups. We introduce Group-Conditional Risk-Averse Calibration (GC-RAC), a framework that extends Risk-Averse Calibration to provide per-group utility certificates for safe decision making. The main idea is to replace a single global calibration price with an aggregated price that combines the contributions of all groups containing a given input, allowing GC-RAC to handle group-specific constraints without relying on nested prediction sets. We develop two variants based on group structure. Mondrian-RAC addresses partition groups, where each input belongs to exactly one group, and provides exact distribution-free per-group guarantees. Multivalid-RAC handles overlapping groups, where an input may belong to several groups, and provides asymptotic simultaneous guarantees using monotone coordinate ascent over group-specific prices. On the UCI Adult dataset across 5 random seeds, Mondrian-RAC reduces the worst-case group coverage violation from 0.025 ± 0.004 under Marginal-RAC to 0.002 ± 0.002, with less than 1% utility loss. On a synthetic task with three overlapping groups, Multivalid-RAC reduces the violation from 0.051 to 0.007 and converges in about two coordinate-ascent iterations. Six ablation studies further show stable performance across calibration set sizes, group imbalance, heterogeneous risk levels, classifier choices, grid resolutions, and utility asymmetry.
  19. Vision-Grounded Kernel Semantic Entropy for Spatial Hallucination Detection in Medical Visual Question Answering
    Adinath Madhavrao Dukre, Yifan Lu, Govinda Kolli, Ziyun Zou, Youssef Mohamed, Azmine Toushik Wasi, Yutong Xie, Shoaib Jameel, Imran Razzak

    Multimodal large language models (MLLMs) for medical Visual Question Answering (VQA) frequently produce spatial hallucination responses referencing incorrect anatomical locations while appearing plausible. Existing detection methods like semantic entropy (SE) and VASE use global perturbations and hard clustering, failing to assess whether models attend to correct regions or capture fine-grained medical semantics (e.g., treating ``right lower lobe'' and ``right middle lobe'' as entirely dissimilar). We propose Vision-Grounded Kernel Semantic Entropy (VG-KSE), which introduces: (1) anatomically-targeted perturbation via a counterfactual framework comparing uncertainty when relevant versus irrelevant regions are masked, yielding a Grounding Inconsistency Score (GIS); and (2) soft semantic kernels with medical-domain embeddings and von Neumann entropy, preserving nuanced terminology relationships. On MIMIC-Diff-VQA and VQA-RAD with CheXagent and LLaVA-Med, VG-KSE consistently outperforms all baselines, achieving +11.7% Spatial-AUC improvement where prior methods perform near chance level.
  20. Position: Biological AI Safety and Regulation Can Learn from Cybersecurity
    Azmine Toushik Wasi, Mahfuz Ahmed Anik, Mst Rafia Islam, Md Manjurul Ahsan, TH Rafi, Dong-Kyu Chae
    ACM FAccT 2027 â–Ș

  21. BanglaHealthMyth: Evaluating LLM Resilience to Health-related Misinformation in Bangladesh
    Noor Mairukh Khan Arnob*, Azmine Toushik Wasi*, Md Noushad Jahan Ramim, Wahid Faisal, Labiba Akram Pritha, Saiyara Mahmud, Md Rizwan Parvez

    Large language models (LLMs) face a critical safety gap in culturally diverse health contexts: standard evaluations assume good-faith queries, while real-world misinformation in low-resource settings is often embedded in elder authority, religious legitimacy, and emotional coercion. We introduce \textsc{BanglaHealthMyth}, a benchmark of 1,015 Bangla health myths, each paired with evidence-based refutations and cultural metadata, designed for adversarial safety testing. Unlike general fact-checking datasets, it is negative-only: all claims are false, and the key metric is \emph{unsafe acceptance}, i.e., when models label false health claims as \textit{True} under adversarial pressure. We evaluate 16 multilingual LLMs using nine culturally grounded adversarial prompt strategies across three manipulation families: authority/legitimacy pressure, pseudo-scientific overconfidence, and emotional/moral coercion. Unsafe acceptance increases by up to 97 percentage points under attack, with myth-rejection rates dropping below 2% for several models. These findings reveal systematic weaknesses in current LLM safety training for culturally embedded, low-resource health misinformation, and highlight the need for culturally aware adversarial evaluation as a core safety requirement.
  22. Generative AI as a Geopolitical Factor in Industry 5.0: Sovereignty, Access, and Control
    Azmine Toushik Wasi, Enjamamul Haque Eram, Sabrina Afroz Mitu, Md Manjurul Ahsan
    ACM Transactions â–Ș [arXiv]

  23. Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset and Benchmarks
    Azmine Toushik Wasi, MD Shafikul Islam Sohan, Adipto Raihan Akib, Mahathik Muhammad Bappy
    ACM Transactions â–Ș [arXiv] â–Ș [GitHub] â–Ș [GitHub] (~22⭐) â–Ș

and more...




*Equal Contribution
$Advising and Mentoring Role
#Corresponding Author






Copyright © 2022-2026 Azmine Toushik Wasi. All rights reserved.