Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Data mining methodological weaknesses and suggested fixes

dc.contributor.authorMaindonald, John
dc.coverage.spatialSydney Australia
dc.date.accessioned2015-12-07T22:47:45Z
dc.date.createdNovember 29-30 2006
dc.date.issued2006
dc.date.updated2016-02-24T10:06:21Z
dc.description.abstractPredictive accuracy claims should give explicit descriptions of the steps followed, with access to the code used. This allows referees and readers to check for common traps, and to repeat the same steps on other data. Feature selection and/or model selection and/or tuning must be independent of the test data. For use of cross-validation, such steps must be repeated at each fold. Even then, such accuracy assessments have the limitation that the target population, to which results will be applied, is commonly different from the source population. Commonly, it is shifted forward in time, and it may differ in other respects also. A consequence of source/target differences is that highly sophisticated modeling may be pointless or even counter-productive. At best, model effects in the target population may be broadly similar. Investigation of the pattern of changes over time is required. Such studies are unusual in the data mining literature, in part because relevant data have not been available. Several recent investigations are noted that shed interesting light on the comparison between observational and experimental studies, with particular relevance when there is an interest in giving parameter estimates a causal interpretation. Data mining activity would benefit from wider co-operation in the development and deployment of computing tools, and from better integration of those tools into the publication process.
dc.identifier.isbn1920682422
dc.identifier.urihttp://hdl.handle.net/1885/26185
dc.publisherAustralian Computer Society Inc.
dc.relation.ispartofseriesAustralasian Data Mining Conference (AusDM 2006)
dc.sourceProceedings of the fifth Australasian Data Mining Conference (AusDM2006)
dc.subjectKeywords: Accuracy assessment; Computing tools; Cross validation; Experimental studies; Mining activities; Model Selection; Observational data; Parameter estimate; Predictive accuracy; Reject inference; Selection bias; Source population; Test data; Inference engine Comparison of algorithms; Data mining; Observational data; Predictive accuracy; Reject inference; Selection bias; Statistics; Target population
dc.titleData mining methodological weaknesses and suggested fixes
dc.typeConference paper
local.bibliographicCitation.lastpage16
local.bibliographicCitation.startpage9
local.contributor.affiliationMaindonald, John, College of Physical and Mathematical Sciences, ANU
local.contributor.authoruidMaindonald, John, u9801539
local.description.embargo2037-12-31
local.description.notesImported from ARIES
local.description.refereedYes
local.identifier.absfor010202 - Biological Mathematics
local.identifier.ariespublicationu3488905xPUB43
local.identifier.scopusID2-s2.0-84870549537
local.type.statusPublished Version

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
01_Maindonald_Data_mining_methodological_2006.pdf
Size:
465.67 KB
Format:
Adobe Portable Document Format