Visual Prompting in LLMs for Enhancing Emotion Recognition
| dc.contributor.author | Zhang, Qixuan | en |
| dc.contributor.author | Wang, Zhifeng | en |
| dc.contributor.author | Zhang, Dylan | en |
| dc.contributor.author | Niu, Wenjia | en |
| dc.contributor.author | Caldwell, Sabrina | en |
| dc.contributor.author | Gedeon, Tom | en |
| dc.contributor.author | Liu, Yang | en |
| dc.contributor.author | Qin, Zhenyue | en |
| dc.date.accessioned | 2025-05-23T03:24:40Z | |
| dc.date.available | 2025-05-23T03:24:40Z | |
| dc.date.issued | 2024 | en |
| dc.description.abstract | Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing. Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped. Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context. To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely. SoV improves accuracy in face count and emotion categorization while preserving the enriched image context. Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments. Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. | en |
| dc.description.status | Peer-reviewed | en |
| dc.format.extent | 16 | en |
| dc.identifier.isbn | 9798891761643 | en |
| dc.identifier.other | ORCID:/0000-0003-0605-3149/work/184099321 | en |
| dc.identifier.scopus | 85217799721 | en |
| dc.identifier.uri | http://www.scopus.com/inward/record.url?scp=85217799721&partnerID=8YFLogxK | en |
| dc.identifier.uri | https://hdl.handle.net/1885/733751101 | |
| dc.language.iso | en | en |
| dc.publisher | Association for Computational Linguistics (ACL) | en |
| dc.relation.ispartof | EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference | en |
| dc.relation.ispartofseries | 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024 | en |
| dc.relation.ispartofseries | EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference | en |
| dc.rights | Publisher Copyright: © 2024 Association for Computational Linguistics. | en |
| dc.title | Visual Prompting in LLMs for Enhancing Emotion Recognition | en |
| dc.type | Conference paper | en |
| dspace.entity.type | Publication | en |
| local.bibliographicCitation.lastpage | 4499 | en |
| local.bibliographicCitation.startpage | 4484 | en |
| local.contributor.affiliation | Zhang, Qixuan; Australian National University | en |
| local.contributor.affiliation | Wang, Zhifeng; Australian National University | en |
| local.contributor.affiliation | Zhang, Dylan; Quriosity Pty Ltd. | en |
| local.contributor.affiliation | Niu, Wenjia; Webumate Pty Ltd. | en |
| local.contributor.affiliation | Caldwell, Sabrina; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.contributor.affiliation | Gedeon, Tom; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.contributor.affiliation | Liu, Yang; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.contributor.affiliation | Qin, Zhenyue; Yale University | en |
| local.identifier.doi | 10.18653/v1/2024.emnlp-main.257 | en |
| local.identifier.pure | a1a97441-53bf-4660-95ff-3d58f9652cba | en |
| local.identifier.url | https://www.scopus.com/pages/publications/85217799721 | en |
| local.type.status | Published | en |