CertifiedGPT: Evaluating Adversarial Robustness of Vision and Language Models Against Targeted Black-Box Attacks

Authors

DOI:

https://doi.org/10.5753/jbcs.2026.6584

Keywords:

Visual Language Models, Fine-tuning, Certified robustness, Randomized smoothing, Adversarial attacks

Abstract

Multimodal Vision and Language Models (VLMs) have attracted significant academic interest due to their strong performance in tasks that require simultaneous inference across image and text modalities. The impact of VLMs is evident both in general society, through powerful chatbots such as ChatGPT and Google Gemini, and in specific industry sectors, including industrial processes, biomedical engineering, systems engineering, and medicine. In contrast to proprietary models, open-source VLMs such as MiniGPT-4 have emerged, offering compact alternatives that demonstrate strong performance. However, such VLMs show certain vulnerabilities to noisy input data (e.g., images), which can compromise the reliability of the output and can be explored by adversarial attacks, even under black-box access. More sophisticated attacks can also force a VLM to generate specific output using targeted attack configurations. In this work, we explored the concept of robustness certification in order to make a VLM robust against targeted black-box attacks through the use of a randomized smoothing technique. First, we fine-tuned the MiniGPT-4 model on a small subset of VQAv2 and applied Gaussian noise to the input images. We also adapted a smoothed method that encapsulates the original decoder and generates the most probable answer, even with the noise in the input images. Finally, we evaluated the model against targeted black-box attacks. Our certified version of MiniGPT-4, when evaluated on a small VQAv2 subset, produced statistically significant results, demonstrating that randomized smoothing is a feasible approach to certifying the robustness of VLMs, especially in scenarios where access to high-performance GPU is limited.

Downloads

Download data is not yet available.

References

Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. (2021). Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3558-3568. DOI: 10.1109/cvpr46437.2021.00356.

Cohen, J., Rosenfeld, E., and Kolter, Z. (2019). Certified adversarial robustness via randomized smoothing. In international conference on machine learning, pages 1310-1320. PMLR. DOI: 10.48550/arxiv.1902.02918.

Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. DOI: 10.48550/arxiv.1810.04805.

Fares, S., Ziu, K., Aremu, T., Durasov, N., Takáč, M., Fua, P., Nandakumar, K., and Laptev, I. (2024). Mirrorcheck: Efficient adversarial defense for vision-language models. arXiv preprint arXiv:2406.09250. DOI: 10.48550/arxiv.2406.09250.

Gan, Z., Chen, Y.-C., Li, L., Zhu, C., Cheng, Y., and Liu, J. (2020). Large-scale adversarial training for vision-and-language representation learning. Advances in Neural Information Processing Systems, 33:6616-6628. DOI: 10.48550/arxiv.2006.06195.

Gao, K., Bai, Y., Bai, J., Yang, Y., and Xia, S.-T. (2024). Adversarial robustness for visual grounding of multimodal large language models. In ICLR 2024 Workshop on Reliable and Responsible Foundation Models. DOI: 10.48550/arxiv.2405.09981.

Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. (2017). Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904-6913. DOI: 10.1007/s11263-018-1116-0.

Islam, S., Elmekki, H., Elsebai, A., Bentahar, J., Drawel, N., Rjoub, G., and Pedrycz, W. (2023). A comprehensive survey on applications of transformers for deep learning tasks. Expert Systems with Applications, page 122666. DOI: 10.1016/j.eswa.2023.122666.

Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25. DOI: 10.1145/3065386.

Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083. DOI: 10.48550/arxiv.1706.06083.

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021). Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748-8763. PmLR. DOI: 10.48550/arxiv.2103.00020.

Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018). Improving language understanding by generative pre-training. Available at:[link].

Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684-10695. DOI: 10.1109/cvpr52688.2022.01042.

Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. (2021). Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114. DOI: 10.48550/arxiv.2111.02114.

Sun, J., Wang, C., Wang, J., Zhang, Y., and Xiao, C. (2024). Safeguarding vision-language models against patched visual prompt injectors. arXiv preprint arXiv:2405.10529. DOI: 10.48550/arxiv.2405.10529.

Uppal, S., Bhagat, S., Hazarika, D., Majumder, N., Poria, S., Zimmermann, R., and Zadeh, A. (2022). Multimodal research in vision and language: A review of current and emerging trends. Information Fusion, 77:149-171. DOI: 10.1016/j.inffus.2021.07.009.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. DOI: 10.65215/r5bs2d54.

Wang, X., Ji, Z., Ma, P., Li, Z., and Wang, S. (2023). Instructta: Instruction-tuned targeted attack for large vision-language models. arXiv preprint arXiv:2312.01886. DOI: 10.1186/s42400-026-00612-4.

Wei, L., Jiang, Z., Huang, W., and Sun, L. (2023). Instructiongpt-4: A 200-instruction paradigm for fine-tuning minigpt-4. arXiv preprint arXiv:2308.12067. Available at:[link].

Yigit, G. and Amasyali, M. F. (2023). From text to multimodal: A comprehensive survey of adversarial example generation in question answering systems. arXiv e-prints, pages arXiv-2312. DOI: 10.48550/arXiv.2312.16156.

Yin, Z., Ye, M., Zhang, T., Du, T., Liu, H., Cheng, J., Wang, T., and Ma, F. (2023). Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models. DOI: 10.52202/075280-2303.

Zhang, A., Lipton, Z. C., Li, M., and Smola, A. J. (2023). Dive into deep learning. Cambridge University Press. DOI: 10.48550/arxiv.2106.11342.

Zhang, J. and Li, C. (2019). Adversarial examples: Opportunities and challenges. IEEE transactions on neural networks and learning systems, 31(7):2578-2593. DOI: 10.1109/tnnls.2019.2933524.

Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.-M. M., and Lin, M. (2023). On evaluating adversarial robustness of large vision-language models. Advances in Neural Information Processing Systems, 36:54111-54138. DOI: 10.52202/075280-2355.

Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. (2023). Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592. DOI: 10.48550/arxiv.2304.10592.

Downloads

Published

2026-08-28

How to Cite

Souza, L., Moura, P. N. de S., & Gatti, M. (2026). CertifiedGPT: Evaluating Adversarial Robustness of Vision and Language Models Against Targeted Black-Box Attacks. Journal of the Brazilian Computer Society, 32(1), 2183–2202. https://doi.org/10.5753/jbcs.2026.6584

Issue

Section

Regular Issue