EXAGREE: Mitigating Explanation Disagreement with Stakeholder-Aligned Models
| dc.contributor.author | Li, Sichao | en |
| dc.contributor.author | Deng, Quanling | en |
| dc.contributor.author | Barnard, Amanda S. | en |
| dc.date.accessioned | 2026-03-20T16:40:36Z | |
| dc.date.available | 2026-03-20T16:40:36Z | |
| dc.date.issued | 2024 | en |
| dc.description.abstract | Conflicting explanations, arising from different attribution methods or model internals, limit the adoption of machine learning models in safety-critical domains. We turn this disagreement into an advantage and introduce EXplanation AGREEment (EXAGREE), a two-stage framework that selects a Stakeholder-Aligned Explanation Model (SAEM) from a set of similar-performing models. The selection maximizes Stakeholder-Machine Agreement (SMA), a single metric that unifies faithfulness and plausibility. EXAGREE couples a differentiable mask-based attribution network (DMAN) with monotone differentiable sorting, enabling gradient-based search inside the constrained model space. Experiments on six real-world datasets demonstrate simultaneous gains of faithfulness, plausibility, and fairness over baselines, while preserving task accuracy. Extensive ablation studies, significance tests, and case studies confirm the robustness and feasibility of the method in practice. | en |
| dc.description.status | Not peer-reviewed | en |
| dc.format.extent | 24 | en |
| dc.identifier.other | dblp:journals/corr/abs-2411-01956 | en |
| dc.identifier.other | ORCID:/0000-0002-6159-1233/work/208815387 | en |
| dc.identifier.uri | https://hdl.handle.net/1885/733807530 | |
| dc.language.iso | en | en |
| dc.rights | DBLP License: DBLP's bibliographic metadata records provided through http://dblp.org/ are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions. | en |
| dc.source | CoRR | en |
| dc.title | EXAGREE: Mitigating Explanation Disagreement with Stakeholder-Aligned Models | en |
| dc.type | Journal article | en |
| dspace.entity.type | Publication | en |
| local.contributor.affiliation | Li, Sichao; ANU College of Systems and Society, The Australian National University | en |
| local.contributor.affiliation | Deng, Quanling; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.contributor.affiliation | Barnard, Amanda S.; School of Computing, ANU College of Systems and Society, The Australian National University | en |
| local.identifier.doi | 10.48550/arXiv.2411.01956 | en |
| local.identifier.pure | 07fdeaf8-f680-4c35-b456-fdcbdc861bc9 | en |
| local.type.status | Published | en |