Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Topic Modelling in Spontaneous Speech Data

dc.contributor.authorReverter-Rambaldi, Marcel
dc.date.accessioned2022-12-08T05:10:38Z
dc.date.available2022-12-08T05:10:38Z
dc.date.issued2022
dc.description.abstractThe development of large-scale, language corpora has highlighted the increasing need for automated methods, to assist humans in the inefficient task of sorting and labelling language-transcripts by semantic contents (i.e. topics). One approach to semantic labelling involves using a class of unsupervised, machine-learning algorithms known as “topic modelling”. These algorithms process a document (e.g. a transcript), and identify clusters representing words that occur in proximity to each other in the document. To date, topic modelling has been implemented widely in written language – including newspapers, academic articles, and business reports – but much less to spontaneous speech data. The linguistics literature has identified the need to apply more qualitative and analytic approaches, when judging and improving topic modelling for future use. My research applies topic-modelling algorithms to transcripts from sociolinguistic interviews, compiled for the Sydney Speaks Project. I apply certain modifications to improve topic-modelling’s performance, including the use of a custom stoplist, a human benchmark for measuring efficacy, and linguistically-based, text partitioning. The findings support the idea that text partitioning and a custom stoplist, produce results that align better with the human benchmark.en_AU
dc.identifier.urihttp://hdl.handle.net/1885/281664
dc.language.isoen_AUen_AU
dc.subjectapplied linguisticsen_AU
dc.subjectcomputational linguisticsen_AU
dc.subjectcorpus linguisticsen_AU
dc.subjectmachine learningen_AU
dc.subjectnatural language processingen_AU
dc.subjecttopic modellingen_AU
dc.subjecttopic modelingen_AU
dc.subjectSydney Speaksen_AU
dc.subjectLanguage Data Commons of Australiaen_AU
dc.subjectLDaCAen_AU
dc.titleTopic Modelling in Spontaneous Speech Dataen_AU
dc.typeThesis (Honours)en_AU
dcterms.valid2022en_AU
local.contributor.affiliationSchool of Literature, Languages and Linguistics, College of Arts and Social Sciences, The Australian National Universityen_AU
local.contributor.supervisorTravis, Catherine
local.description.notesthe author deposited 8.12.2022en_AU
local.identifier.doi10.25911/M1YF-ZF55
local.mintdoiminten_AU
local.type.degreeOtheren_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Marcel_Honours_Thesis_Publishable.pdf
Size:
1.44 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
884 B
Format:
Item-specific license agreed upon to submission
Description: