Human Perceptions Meet AI Analysis: A Usability Evaluation of Climate Conference Websites
DOI:
https://doi.org/10.5753/jis.2026.7129Keywords:
Generative AI, Usability Evaluation, Human-Computer Interaction, Climate Conference WebsitesAbstract
Introduction: Digital platforms play a central role in mediating major international events, serving as gateways for information, mobilization, and public engagement. Evaluating usability in heterogeneous, cross-cultural contexts such as the Conference of the Parties (COP) poses significant challenges. The emergence of Generative Artificial Intelligence (GAI) offers new opportunities for automated usability assessment, but its alignment with real user experience remains underexplored. Objective: This study critically compares GAI-generated usability evaluations against human perceptions of climate conference websites. GAIs assessed both COP29 and COP30 platforms, while human evaluations focused exclusively on COP30 during the event in Belém, Brazil. The objective is to examine convergences and divergences between automated and human-centered assessments. Methods: Four GAIs (ChatGPT, Gemini, DeepSeek, and Copilot) performed heuristic evaluations based on ISO 9241-210 criteria. GAI outputs were analyzed using Directed Categorical Content Analysis. These findings were complemented by an exploratory evaluation involving ten human participants (n=10) with diverse profiles, whose perceptions were collected via an online questionnaire during COP30. Final analysis involved triangulation of GAI findings with human feedback. Results: Results revealed a notable divergence in overall assessment. Humans reported high satisfaction and usability for critical tasks, contrasting with GAIs, which penalized COP30 for structural deficiencies, particularly low Controllability. However, convergence occurred in Individualization Adequacy, the lowest-scoring criterion for both GAIs and humans, confirming systemic failures in adaptability and clarity. Conclusion: GAIs and users converge on structural issues but diverge on subjective and contextual evaluations. This highlights the value of hybrid models: GAIs provide systematic heuristic data, while human feedback grounds practical and sociocultural interpretations. Future work should include longitudinal evaluations, automated accessibility checkers, and hybrid workflows for robust assessment frameworks.
Downloads
References
Alsadi, A. and Miller, D. (2023). Exploring the impact of artificial intelligence language model chatgpt on the user experience. International Journal of Technology, Innovation and Management (IJTIM), 1:2023. DOI: https://doi.org/10.54489/ijtim.v3i1.195.
Araújo, M. F. B. d., Mota, M. P., and Seruffo, M. C. d. R. (2024). Combinando inteligência artificial generativa e inspeção humana: Uma análise da usabilidade do site da greenpeace brasil. In Anais Estendidos do XXIII Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais (IHC), pages 81–85. Sociedade Brasileira de Computação. DOI: https://doi.org/10.5753/ihc_estendido.2024.243962.
Bangor, A., Kortum, P. T., and Miller, J. T. (2008). An empirical evaluation of the system usability scale. Intl. Journal of Human–Computer Interaction, 24(6):574–594. DOI: https://doi.org/10.1080/10447310802205776.
Bardin, L. (2011). Análise de conteúdo. ed. Revista e Ampliada. São Paulo: Edições, 70.
Bevan, N., Carter, J., and Harker, S. (2015). Iso 9241-11 revised: What have we learnt about usability since 1998? In Human-Computer Interaction: Design and Evaluation: 17th International Conference, HCI International 2015, Los Angeles, CA, USA, August 2-7, 2015, Proceedings, Part I 17, pages 143–151. Springer. DOI: https://doi.org/10.1007/978-3-319-20901-2_13.
Bleichner, A. and Hermansson, N. (2023). Investigating the usefulness of a generative ai when designing user interfaces. Master's thesis, Uppsala University.
Borges, J. M. and Araújo, R. D. (2024). Experiences and challenges of a redesign process with the support of an ai assistant on an educational platform. New York, NY, USA. Association for Computing Machinery. DOI: https://doi.org/10.1145/3702038.3702046.
Braun, V. and Clarke, V. (2006). Using thematic analysis in psychology. Qualitative research in psychology, 3(2):77–101. DOI: https://doi.org/10.1191/1478088706qp063oa.
Brazilian Federal Government (2024). COP 30 in Brazil. [link]. Accessed: 22 August 2026.
Brooke, J. (1996). Sus: A 'quick and dirty' usability scale. In Jordan, P. W., Thomas, B., Weerdmeester, B. A., and McClelland, I. L., editors, Usability Evaluation in Industry, pages 189–194. Taylor & Francis, London.
Campos, T., Castello, M., Damasceno, E., and Valentim, N. (2025). An updated systematic mapping study on usability and user experience evaluation of touchable holographic solutions. Journal on Interactive Systems, 16(1):172–198. DOI: https://doi.org/10.5753/jis.2025.4694.
Carifio, J. and Perla, R. J. (2007). Ten common misunderstandings, misconceptions, persistent myths and urban legends about likert scales and likert response formats and their antidotes. Journal of social sciences, 3(3):106–116. DOI: http://dx.doi.org/10.3844/jssp.2007.106.116.
Castells, M. (2008). The new public sphere: Global civil society, communication networks, and global governance. The Annals of the American Academy of Political and Social Science, 616(1):78–93. DOI: https://doi.org/10.1177/0002716207311877.
Cyr, D., Head, M., and Ivanov, A. (2006). Design aesthetics leading to m-loyalty in mobile commerce. Information & management, 43(8):950–963. DOI: https://doi.org/10.1016/j.im.2006.08.009.
de Carvalho, J. F. M., Moura, F. R. T., do Carmo Pereira, L. G., da Conceição Estevam, L., Negreiros, W. J. A. G., Alves, A. V. N., Pinto, L. V. L., dos Santos, A. J. S., Júnior, W. d. S. O., Cardoso, D. L., et al. (2026). Innovations in geometry teaching with miritiboard vr 2.0 glasses: Customizations and usability evaluation. Journal on Interactive Systems, 17(1):175–192. DOI: https://doi.org/10.5753/jis.2026.5755.
de Oliveira, L., Amaral, M., Bim, S., Valença, G., Almeida, L., Salgado, L., Gasparini, I., and da Silva, C. (2024). Grandihc-br 2025-2035 – gc3: Plurality and decoloniality in hci. In Anais do XXIII Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 1067–1085, Porto Alegre, RS, Brasil. SBC. DOI: https://doi.org/10.1145/3702038.3702056.
Duarte, E. F., Toledo Palomino, P., Pontual Falcão, T., Porto, G. L. P. M. B., Portela, C. d. S., Ribeiro, D. F., Nascimento, A., Costa Aguiar, Y. P., Souza, M., Moutin Segoria Gasparotto, A., et al. (2024). Grandihc-br 2025-2035-gc6: Implications of artificial intelligence in hci: A discussion on paradigms ethics and diversity equity and inclusion. In Proceedings of the XXIII Brazilian Symposium on Human Factors in Computing Systems, pages 1–19. DOI: https://doi.org/10.1145/3702038.3702059.
Faulkner, L. (2003). Beyond the five-user assumption: Benefits of increased sample sizes in usability testing. Behavior Research Methods, Instruments, & Computers, 35(3):379–383. DOI: https://doi.org/10.3758/BF03195514.
Fischer, M. and Lanquillon, C. (2024). Evaluation of generative ai-assisted software design and engineering: A user-centered approach. In Degen, H. and Ntoa, S., editors, Artificial Intelligence in HCI, pages 31–47, Cham. Springer Nature Switzerland. DOI: https://dl.acm.org/doi/10.1007/978-3-031-60606-9_3.
Hassenzahl, M. and Tractinsky, N. (2006). User experience-a research agenda. Behaviour & information technology, 25(2):91–97. DOI: https://doi.org/10.1080/01449290500330331.
Hofstede, G. (2001). Culture's Consequences: Comparing Values, Behaviors, Institutions, and Organizations Across Nations. Sage Publications, Thousand Oaks, 2 edition.
Hornbæk, K. and Hertzum, M. (2017). Technology acceptance and user experience: A review of the experiential component in hci. ACM Transactions on Computer-Human Interaction (TOCHI), 24(5):1–30. DOI: https://doi.org/10.1145/3127358.
International Organization for Standardization (2019). ISO 9241-210: Ergonomics of human-system interaction – Part 210: Human-centred design for interactive systems. [link]. Accessed: 31 August 2026.
Ivory, M. Y. and Hearst, M. A. (2001). The state of the art in automating usability evaluation of user interfaces. ACM Comput. Surv., 33(4):470–516. DOI: https://doi.org/10.1145/503112.503114.
Kuric, E., Demcak, P., Krajcovic, M., and Lang, J. (2025). Systematic literature review of automation and artificial intelligence in usability issue detection. DOI: https://doi.org/10.48550/arXiv.2504.01415.
Lewis, J. R. (2018). The system usability scale: past, present, and future. International Journal of Human–Computer Interaction, 34(7):577–590. DOI: https://doi.org/10.1080/10447318.2018.1455307.
Li, D. (2024). Exploration of user experience design optimization for the campus library information management system. Journal of Education, Humanities and Social Sciences, 37:6–15. DOI: https://doi.org/10.54097/8vpy7360.
Lima, D., Medeiros, A., de Cássia Paulino, R., and Seruffo, M. C. (2025). Can ai judge usability? a comparative analysis of generative tools on climate conference websites. In Anais do XXIV Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 497–517, Porto Alegre, RS, Brasil. SBC. DOI: https://doi.org/10.5753/ihc.2025.10805.
Lima, J. and de Lima, T. C. (2025). Blueprint para design participativo sensível ao contexto: Resposta aos desafios gc3 e gc4 do grandihc-br. In Anais Estendidos do XXIV Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 391–396, Porto Alegre, RS, Brasil. SBC. DOI: https://doi.org/10.5753/ihc_estendido.2025.16215.
Lincoln, Y. (1980). Guba. e.(1985). naturalistic inquiry. Beverly Hills: Sage. LincolnNaturalistic Inquiry1985. DOI: https://doi.org/10.1016/0147-1767(85)90062-8.
Liu, F. (2021). International usability testing: Why you need it. [link]. Accessed: 23 August 2026.
Marcus, A. and Gould, E. W. (2000). Crosscurrents: cultural dimensions and global web userinterface design. Interactions, 7(4):32–46. DOI: https://doi.org/10.1145/345190.345238.
Meskó, B. and Topol, E. J. (2023). The imperative for regulatory oversight of large language models (or generative ai) in healthcare. NPJ digital medicine, 6(1):120. DOI: https://doi.org/10.1038/s41746-023-00873-0.
Miraz, M. H., Ali, M., and Excell, P. S. (2017). Multilingual website usability analysis based on an international user survey. CoRR, abs/1708.05085. DOI: https://doi.org/10.48550/arXiv.1708.05085.
Nielsen, J. (1994a). Estimating the number of subjects needed for a thinking aloud test. International journal of human-computer studies, 41(3):385–397. DOI: https://doi.org/10.1006/ijhc.1994.1065.
Nielsen, J. (1994b). Heuristic evaluation. In Nielsen, J. and Mack, R. L., editors, Usability Inspection Methods, pages 25–62. John Wiley & Sons, New York. DOI: https://dl.acm.org/doi/10.5555/189200.189209.
Nielsen, J. (1999). Designing web usability: The practice of simplicity. New riders publishing.
Nisbett, R. E. (2003). The Geography of Thought: How Asians and Westerners Think Differently... and Why. The Free Press, New York.
Norman Donald, A. (2013). The design of everyday things. MIT Press.
Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D., Breckler, S. J., Buck, S., Chambers, C. D., Chin, G., Christensen, G., et al. (2015). Promoting an open research culture. Science, 348(6242):1422–1425. DOI: https://doi.org/10.1126/science.aab2374.
Nowell, L. S., Norris, J. M., White, D. E., and Moules, N. J. (2017). Thematic analysis: Striving to meet the trustworthiness criteria. International journal of qualitative methods, 16(1):1609406917733847. DOI: https://doi.org/10.1177/1609406917733847.
Oliveira, L. F. P. d. and Ferreira, S. B. L. (2022). Usabilidade e acessibilidade: Um estudo de caso com a plataforma instagram. [link]. Trabalho de Conclusão de Curso, Universidade Federal do Estado do Rio de Janeiro (UNIRIO).
Sandelowski, M. (2000). Whatever happened to qualitative description? Research in nursing & health, 23(4):334–340. DOI: https://doi.org/10.1002/1098-240x(200008)23:4.
Sauro, J. and Lewis, J. R. (2011). When designing usability questionnaires, does it hurt to be positive? In Proceedings of the SIGCHI conference on human factors in computing systems, pages 2215–2224. DOI: https://doi.org/10.1145/1978942.1979266.
Sen, A. (2000). Desenvolvimento como Liberdade. Companhia das Letras, São Paulo.
Serra, L., Carvalho, L., Ferreira, L., Vaz, J., and Freire, A. (2015). Accessibility evaluation of e-government mobile applications in brazil. Procedia Computer Science, 67:348–357. DOI: https://doi.org/10.1016/j.procs.2015.09.279.
Shneiderman, B. and Plaisant, C. (2010). Designing the user interface: strategies for effective human-computer interaction. DOI: https://doi.org/10.1145/25065.950626.
Streiner, D., Norman, G. R., and Cairney, J. (2016). Health measurement scales: a practical guide to their development and use. Aust NZJ Public Health, 40(3):294–5. DOI: https://doi.org/10.1093/med/9780199685219.001.0001.
Tractinsky, N. (2018). The usability construct: a dead end? Human–Computer Interaction, 33(2):131–177. DOI: https://doi.org/10.1080/07370024.2017.1298038.
Vatrapu, R. and Pérez-Quiñones, M. A. (2004). Culture and international usability testing: The effects of culture in structured interviews. CoRR, cs.HC/0405045. DOI: https://doi.org/10.48550/arXiv.cs/0405045.
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382. DOI: https://doi.org/10.48550/arXiv.2302.11382.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Danilo T. Lima, Alana M. Medeiros, Rita de Cássia R. Paulino, Marcos César R. Seruffo

This work is licensed under a Creative Commons Attribution 4.0 International License.
JIS is free of charge for authors and readers, and all papers published by JIS follow the Creative Commons Attribution 4.0 International (CC BY 4.0) license.


