中大學術數位典藏-NCU Institutional Repository:Item 987654321/106900
English  |  正體中文  |  简体中文  |  全文笔数/总笔数 : 94459/94459 (100%)
造访人次 : 87666460      在线人数 : 161
RC Version 7.0 © Powered By DSPACE, MIT. Enhanced by NTU Library IR team.
搜寻范围 查询小技巧:
  • 您可在西文检索词汇前后加上"双引号",以获取较精准的检索结果
  • 若欲以作者姓名搜寻,建议至进阶搜寻限定作者字段,可获得较完整数据
  • 进阶搜寻


    jsp.display-item.identifier=請使用永久網址來引用或連結此文件: https://ir.lib.ncu.edu.tw/handle/987654321/106900


    题名: On mining incomplete medical datasets: Ordering imputation and classification
    作者: 胡雅涵;Chen, Chih-Wen;Lin, Wei-Chao;Ke, Shih-Wen;Tsai, Chih-Fong;Hu, Ya-Han
    贡献者: 管理學院資訊管理學系
    关键词: Algorithms;Data Accuracy;Data Interpretation, Statistical;Data Mining - methods;Data Mining - standards;Humans;Support Vector Machine
    日期: 2015-01-01
    上传时间: 2026-04-23 13:48:13 (UTC+8)
    出版者: IOS Press;London, England: SAGE Publications
    摘要: 摘要: BACKGROUND: To collect medical datasets, it is usually the case that a number of data samples contain some missing values. Performing the data mining task over the incomplete datasets is a difficult problem. In general, missing value imputation can be approached, which aims at providing estimations for missing values by reasoning from the observed data. Consequently, the effectiveness of missing value imputation is heavily dependent on the observed data (or complete data) in the incomplete datasets. OBJECTIVE: In this paper, the research objective is to perform instance selection to filter out some noisy data (or outliers) from a given(complete) dataset to see its effect on the final imputation result. Specifically, four different processes of combining instance selection and missing value imputation are proposed and compared in terms of data classification. METHODS: Experiments are conducted based on 11 medical related datasets containing categorical, numerical, and mixed attribute types of data. In addition, missing values for each dataset are introduced into all attributes (the missing data rates are 10%, 20%, 30%, 40%, and 50%). For instance selection and missing value imputation, the DROP3 and k-nearest neighbor imputation methods are employed. On the other hand, the support vector machine (SVM) classifier is used to assess the final classification accuracy of the four different processes. RESULTS: The experimental results show that the second process by performing instance selection first and imputation second allows the SVM classifiers to outperform the other processes. CONCLUSIONS: For incomplete medical datasets containing some missing values, it is necessary to perform missing value imputation. In this paper, we demonstrate that instance selection can be used to filter out some noisy data or outliers before the imputation process. In other words, the observed data for missing value imputation may contain some noisy information, which can degrade the quality of the imputation result as well as the classification performance.
    其他題名: Technol Health Care
    出版者: London, England: SAGE Publications
    出版日期: 2015-01-01
    出處: Technology and health care, 2015-01, Vol.23 (5), p.619-625
    資源來源: EBSCOhost Academic Search Premier
    版權: IOS Press and the authors. All rights reserved
    識別號: ISSN: 0928-7329
    識別號: ISSN: 1878-7401
    識別號: EISSN: 1878-7401
    識別號: DOI: 10.3233/THC-151018
    識別號: PMID: 26410122
    显示于类别:[資訊管理學系] 期刊論文

    文件中的档案:

    档案 描述 大小格式浏览次数
    index.html0KbHTML23检视/开启


    在NCUIR中所有的数据项都受到原著作权保护.

    社群 sharing

    ::: Copyright National Central University. | 國立中央大學圖書館版權所有 | 收藏本站 | 設為首頁 | 最佳瀏覽畫面: 1024*768 | 建站日期:8-24-2009 :::
    DSpace Software Copyright © 2002-2004  MIT &  Hewlett-Packard  /   Enhanced by   NTU Library IR team Copyright ©   - 隱私權政策聲明