Table 2

Properties of the experiments on streaming feature selection.

AlgorithmProperties
Grafting[38]
  • Single or group feature selection: single.

  • Compared with which algorithms: none.

  • Datasets: Two synthetic datasets (A and B) and Pima Indian Diabetes dataset (Blake & Merz, 1998) [69].

  • Classifiers: Combination of the speed of filters and the accuracy of the wrapper.

  • Environment: Not mentioned.

Alpha investing [40]
  • Single or group feature selection: single.

  • Compared with which algorithms: none. The appraisal was limited to the accuracy of the whole dataset.

  • Datasets: Seven datasets from the UCI [57] repository: cleve, internet, ionosphere, spam, spect, wdbc, and wpbc. Three datasets on gene expression: aml, ha, and hung.

  • Classifiers: C4.5, fivefold cross-validation.

  • Environment: Not mentioned.

OSFS and Fast-OSFS [39]
  • Single or group feature selection: single.

  • Compared with which algorithms: Grafting and alpha investing [71].

  • Datasets: Ten public challenge datasets: lymphoma, ovarian-cancer, breast-cancer, hiva, nova, manelon, arcene, dexter, dorohthea and sido0.

  • Classifiers: k-nn, decision tree (J48) and random forest (Spider 2010).

  • Environment: Windows XP, a 2.6GHz CPU, and 2 GB memory.

SAOLA [44]
  • Single or group feature selection: single.

  • Compared with which algorithms: Fast-OSFS [43], alpha investing [71], OFS [72], FCBF [3], as well as two state-of-the-art algorithms, SPSF-LAR [73] and GDM [74].

  • Datasets: Ten high-dimensional datasets: two public microarray datasets (lung cancer and leukemia), two text-categorization datasets (ohsumed and apcj etiology), two biomedical datasets (hiva and breast cancer), three NIPS 2003 (dexter, madelon, and dorothea) and the thrombin dataset, which was chosen from KDD Cup 2001. Four extremely high-dimensional datasets from the Libsvm dataset website: news20, url1, webspam, and kdd2010.

  • Classifiers: KNN and J48, which are provided in the Spider Toolbox2 [75].

  • Environment: Intel i7-2600 with a 3.4GHz CPU and 24 GB of memory.

OS-NRRSAR-SA [41]
  • Single or group feature selection: single.

  • Compared with which algorithms: Grafting, information investing [71], fast-OSFS, and DIA-RED.

  • Datasets: Fourteen high-dimensional datasets: The dorothea, arcene, dexter, and madelon datasets from the NIPS 2003 Feature-Selection Challenge. The nova, sylva, and hiva datasets from the WCCI 2006 Performance Prediction Challenges. The sido0 and cina0 datasets from the WCCI 2008 Causation and Prediction Challenges. The arrhythmia and multiple features datasets from the UCI Machine Learning Repository. Three synthetic datasets: tm1, tm2, and tm3.

  • Classifiers: J48, JRip, Naive Bayes, and kernel SVM with the RBF kernel function.

  • Environment: Dell workstation with Windows 7, 2GB of memory, and a 2.4 GHz CPU.

DIA-RED [45]
  • Single or group feature selection: single.

  • Compared with which algorithms: None.

  • Datasets: Six datasets from the UCI [57] Machine-Learning Repository: Backup-large, Dermatology, Splice, Kr-vs-kp, Mushroom, and Ticdata2000.

  • Classifiers: information entropy used to measure the uncertainty of a dataset: complementary entropy [76], combination entropy [77], and Shannon’s entropy [78].

  • Environment: Windows 7, an Intel Core i7-2600 CPU (2.66GHz), and 4 GB of memory.

GFSSF [48]
  • Single or group feature selection: singleand Group selection.

  • Compared with which algorithms: Five standard feature-selection algorithms: MIFS [13], joint mutual information [79], mRMR [8], ReliefF [20], and lasso [28]. Four streaming-feature-selection algorithms: grafting [38], α investing [40], OSFS [39], and Fast-OSFS [39]. One group-feature-selection algorithm: group lasso [35].

  • Datasets: Five UCI [57] benchmark datasets: WDBC, WPBC, IONOSPHERE, SPECTF, and ARRHYTHMIA. Five challenge datasets with relatively high feature dimensions) downloaded from http://mldata.org/repository): DLBCL (7,130 features; 77 instances), LUNG (7,130 features; 96 instances), CNS (7,130 features; 96 instances), ARCENE (10,000 features; 100 instances), and OVARIAN (15,155 features; 253 instances). Five UCI [57] datasets with generated group structures: HILL-VALLEY (400 features; 606 instances), NORTHIX (800 features; 115 instances), MADELON (2,000 features; 4,400 instances), ISOLET (2,468 features; 7,797 instances), and MULTI-FEATURES (2,567 features; 2,000 instances).

  • Classifiers: NaiveBayes [80], k-NN [81], C4.5 [82], and Randomforest [83].

  • Environment: Windows 7, a 3.33GHz dual-core CPU, and 4 GB of memory.

group-SAOLA [49]
  • Single or group feature selection: group

  • Compared with which algorithms: Three state-of-the-art online-feature-selection methods:

  • Fast-OSFS [43], alpha investing [40], and OFS [43]. Three batch methods: one well-established algorithm (FCBF) [3], and two state-of-the-art algorithms (SPSF-LAR [73] and GDM [74]).

  • Datasets: Ten high-dimensional datasets: madelon, hiva, leukemia, lung-cancer, ohsumed, breast-cancer, dexter, apcj-etiology, dorothea, and thrombin. Four extremely high-dimensional datasets: news20, url1, webspam, and kdd2010.

  • Classifiers: KNN and J48, which are provided in the Spider Toolbox [75], and SVM.

  • Environment: Intel i7-2600, a 3.4GHz CPU, and 24 GB of memory.

OGFS [50 51]
  • Single or group feature selection: singleand group.

  • Compared with which algorithms: Grafting, alpha investing, and OSFS.

  • Datasets: Eight datasets from UCI: Wdbc, Ionosphere, Spectf, Spambase, Colon, Prostate, Leukemia and Lungcancer. Three datasets from the real world: Soccer, Flower-17, and 15 Scenes.

  • Classifiers: appraisal was based on number of the selected features.

  • Environment: Windows XP, a 2.5GHz CPU, and 2 GB of memory.

or Create an Account

Close subscription notice
Close access options