English  |  正體中文  |  简体中文  |  全文筆數/總筆數 : 94459/94459 (100%)
造訪人次 : 87654678      線上人數 : 297
RC Version 7.0 © Powered By DSPACE, MIT. Enhanced by NTU Library IR team.
搜尋範圍 查詢小技巧:
  • 您可在西文檢索詞彙前後加上"雙引號",以獲取較精準的檢索結果
  • 若欲以作者姓名搜尋,建議至進階搜尋限定作者欄位,可獲得較完整資料
  • 進階搜尋


    請使用永久網址來引用或連結此文件: https://ir.lib.ncu.edu.tw/handle/987654321/106382


    題名: An efficient data preprocessing approach for large scale medical data mining
    作者: 胡雅涵;Hu, Ya-Han;Lin, Wei-Chao;Tsai, Chih-Fong;Ke, Shih-Wen;Chen, Chih-Wen
    貢獻者: 管理學院資訊管理學系
    關鍵詞: Algorithms;Classification;Computational efficiency;Data mining;Data Mining - methods;Datasets as Topic;Decision Trees;Humans;Machine Learning;Mathematical models;Medical;Models, Theoretical;Preprocessing;Training
    日期: 2015-01-01
    上傳時間: 2026-04-23 13:19:24 (UTC+8)
    出版者: IOS Press;London, England: SAGE Publications
    摘要: 摘要: Background: The size of medical datasets is usually very large, which directly affects the computational cost of the data mining process. Instance selection is a data preprocessing step in the knowledge discovery process, which can be employed to reduce storage requirements while also maintaining the mining quality. This process aims to filter out outliers (or noisy data) from a given (training) dataset. However, when the dataset is very large in size, more time is required to accomplish the instance selection task. Objective: In this paper, we introduce an efficient data preprocessing approach (EDP), which is composed of two steps. The first step is based on training a model over a small amount of training data after preforming instance selection. The model is then used to identify the rest of the large amount of training data. Methods: Experiments are conducted based on two medical datasets for breast cancer and protein homology prediction problems that contain over 100000 data samples. In addition, three well-known instance selection algorithms are used, IB3, DROP3, and genetic algorithms. On the other hand, three popular classification techniques are used to construct the learning models for comparison, namely the CART decision tree, k-nearest neighbor (k-NN), and support vector machine (SVM). Results: The results show that our proposed approach not only reduces the computational cost by nearly a factor of two or three over three other state-of-the-art algorithms, but also maintains the final classification accuracy. Conclusions: To perform instance selection over large scale medical datasets, it requires a large computational cost to directly execute existing instance selection algorithms. Our proposed EDP approach solves this problem by training a learning model to recognize good and noisy data. To consider both computational complexity and final classification accuracy, the proposed EDP has been demonstrated its efficiency and effectiveness in the large scale instance selection problem.
    其他題名: Technol Health Care
    出版者: London, England: SAGE Publications
    出版日期: 2015-01-01
    出處: Technology and health care, 2015-01, Vol.23 (2), p.153-160
    資源來源: EBSCOhost Academic Search Premier
    版權: 2015 ‒ IOS Press and the authors. All rights reserved
    識別號: ISSN: 0928-7329
    識別號: ISSN: 1878-7401
    識別號: EISSN: 1878-7401
    識別號: DOI: 10.3233/THC-140887
    識別號: PMID: 25515050
    顯示於類別:[資訊管理學系] 期刊論文

    文件中的檔案:

    檔案 描述 大小格式瀏覽次數
    index.html0KbHTML28檢視/開啟


    在NCUIR中所有的資料項目都受到原著作權保護.

    社群 sharing

    ::: Copyright National Central University. | 國立中央大學圖書館版權所有 | 收藏本站 | 設為首頁 | 最佳瀏覽畫面: 1024*768 | 建站日期:8-24-2009 :::
    DSpace Software Copyright © 2002-2004  MIT &  Hewlett-Packard  /   Enhanced by   NTU Library IR team Copyright ©   - 隱私權政策聲明