Topic Modelling in Spontaneous Speech Data
| dc.contributor.author | Reverter-Rambaldi, Marcel | |
| dc.date.accessioned | 2022-12-08T05:10:38Z | |
| dc.date.available | 2022-12-08T05:10:38Z | |
| dc.date.issued | 2022 | |
| dc.description.abstract | The development of large-scale, language corpora has highlighted the increasing need for automated methods, to assist humans in the inefficient task of sorting and labelling language-transcripts by semantic contents (i.e. topics). One approach to semantic labelling involves using a class of unsupervised, machine-learning algorithms known as “topic modelling”. These algorithms process a document (e.g. a transcript), and identify clusters representing words that occur in proximity to each other in the document. To date, topic modelling has been implemented widely in written language – including newspapers, academic articles, and business reports – but much less to spontaneous speech data. The linguistics literature has identified the need to apply more qualitative and analytic approaches, when judging and improving topic modelling for future use. My research applies topic-modelling algorithms to transcripts from sociolinguistic interviews, compiled for the Sydney Speaks Project. I apply certain modifications to improve topic-modelling’s performance, including the use of a custom stoplist, a human benchmark for measuring efficacy, and linguistically-based, text partitioning. The findings support the idea that text partitioning and a custom stoplist, produce results that align better with the human benchmark. | en_AU |
| dc.identifier.uri | http://hdl.handle.net/1885/281664 | |
| dc.language.iso | en_AU | en_AU |
| dc.subject | applied linguistics | en_AU |
| dc.subject | computational linguistics | en_AU |
| dc.subject | corpus linguistics | en_AU |
| dc.subject | machine learning | en_AU |
| dc.subject | natural language processing | en_AU |
| dc.subject | topic modelling | en_AU |
| dc.subject | topic modeling | en_AU |
| dc.subject | Sydney Speaks | en_AU |
| dc.subject | Language Data Commons of Australia | en_AU |
| dc.subject | LDaCA | en_AU |
| dc.title | Topic Modelling in Spontaneous Speech Data | en_AU |
| dc.type | Thesis (Honours) | en_AU |
| dcterms.valid | 2022 | en_AU |
| local.contributor.affiliation | School of Literature, Languages and Linguistics, College of Arts and Social Sciences, The Australian National University | en_AU |
| local.contributor.supervisor | Travis, Catherine | |
| local.description.notes | the author deposited 8.12.2022 | en_AU |
| local.identifier.doi | 10.25911/M1YF-ZF55 | |
| local.mintdoi | mint | en_AU |
| local.type.degree | Other | en_AU |