Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

How to tell Real From Fake? Understanding how to classify human-authored and machine-generated text

dc.contributor.authorDebashish, Chakraborty
dc.date.accessioned2023-05-01T06:49:15Z
dc.date.available2023-05-01T06:49:15Z
dc.date.issued2019
dc.description.abstractNatural Language Generation (NLG) using Generative Adversarial Networks (GANs) has been an active field of research as it alleviates restrictions in conventional Language Modelling based text generators e.g. Long-Short Term Memory (LSTM) networks. The adequacy of a GAN-based text generator depends on its capacity to classify human-written (real) and machine-generated (synthetic) text. However, traditional evaluation metrics used by these generators cannot effectively capture classification features in NLG tasks, such as creative writing. We prove this by using an LSTM network to almost perfectly classify sentences generated by a LeakGAN, a state-of-the-art GAN for long text generation. This thesis attempts a rare approach to understand real and synthetic sentences using meaningful and interpretable features of long sentences (with at least 20 words). We analyse novelty and diversity features of real and synthetic sentences, generate by a LeakGAN, using three meaningful text dissimilarity functions: Jaccard Distance (JD), Normalised Levenshtein Distance (NLD) and Word Mover’s Distance (WMD). In particular, these functions focus on (1) the number of common words, (2) the order of these words, and (3) the semantic similarity in both sentence types, making them interpretable. We provide a comprehensive investigation to identify the effectiveness of novelty and diversity, in classifying real and synthetic sentences, by training two different classification algorithms of varying complexities. Our evaluations show that sentence diversities, using JD and NLD, are the most effective features for classification of human-authored and machine-generated sentences.en_AU
dc.identifier.urihttp://hdl.handle.net/1885/289806
dc.language.isoen_AUen_AU
dc.subjectMachine Learningen_AU
dc.subjectGenerative Artificial Intelligenceen_AU
dc.subjectArtificial Intelligenceen_AU
dc.subjectNatural Language Processingen_AU
dc.titleHow to tell Real From Fake? Understanding how to classify human-authored and machine-generated texten_AU
dc.typeThesis (Masters)en_AU
dcterms.valid2019en_AU
local.contributor.affiliationANU College of Engineering, Computing & Cybernetics, The Australian National Universityen_AU
local.contributor.supervisorHaslum, Patrik
local.description.notesthe author deposited 1 May 2023en_AU
local.identifier.doi10.25911/81VN-SA22
local.mintdoiminten_AU
local.type.degreeOtheren_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Debashish_Chakraborty_thesis.pdf
Size:
1.8 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
884 B
Format:
Item-specific license agreed upon to submission
Description: