Exploring the stability of feature selection methods across a palette of gene expression datasets

Mungloo-Dilmohamud, Zahra; Jaufeerally-Fakim, Yasmina; Peña-Reyes, Carlos Andrés

doi:10.1145/3375923.3375938

Exploring the stability of feature selection methods across a palette of gene expression datasets

Mungloo-Dilmohamud, Zahra; Jaufeerally-Fakim, Yasmina; Peña-Reyes, Carlos Andrés

2019

Herunterladen

Formate

Formate
BibTeX
MARCXML
TextMARC
MARC
DublinCore
EndNote
NLM
RefWorks
RIS

Cite

Résumé

Gene expression data often need to be classified into classes or grouped into clusters for further analysis, using different machine learning techniques and an important pre-processing step is feature selection (FS). The aim of this study is to investigate the stability of some diverse FS methods on a plethora of microarray gene expression data. This experimental work is broken into three parts. Step 1 involves running some FS methods on one gene expression dataset to have a preliminary assessment on the similarity, or dissimilarity, of the resulting feature subsets across methods. Step 2 involves running two of these methods on a large number of different datasets to investigate whether the results produced by the methods are dependent on the features of the dataset: binary, multiclass, small or large dataset. The final step explores how the similarity of selected feature subsets between pairs of methods evolves as the size of the subsets are increased. Results show that the studied methods display a high amount of variability in terms of the resulting selected features. The feature subsets differed both inter- and intra- methods for different datasets. The reason behind this is not clear yet and is being further investigated. The final objective of the research, that is to define how to select a FS method, is an ongoing work whose initial findings are reported herein.

Einzelheiten

Titel

Exploring the stability of feature selection methods across a palette of gene expression datasets

Autor(en)/ in(nen)

Mungloo-Dilmohamud, Zahra (Department of DT, FoICDT, University of Mauritius, Reduit, Mauritius)
Jaufeerally-Fakim, Yasmina (Biotechnology Dept., FoA, University of Mauritius, Reduit, Mauritius)
Peña-Reyes, Carlos Andrés (School of Engineering and Management Vaud, HES-SO, University of Applied Sciences and Arts Western Switzerland)

Datum

2019-11

Veröffentlich in

ICBBE '19: Proceedings of the 2019 6th International Conference on Biomedical and Bioinformatics Engineering, 13-15 November 2019, Shanghai, China

Verlag

Shanghai, China, 13-15 November 2019

Seitenzahl & Äquivalente

6 p.

Vorgestellt auf

6th International Conference on Biomedical and Bioinformatics Engineering, Shanghai, China, 2019-11-13, 2019-11-15

ISBN

9781450372992

DOI

https://doi.org/10.1145/3375923.3375938

Schlüsselwörter

applied computing ; life and medical sciences ; bioinformatics

Papiertyp

full paper

Domaine

Ingénierie et Architecture

Ecole

HEIG-VD

Institut

IICT - Institut des Technologies de l'Information et de la Communication

Das Dokument erscheint in

Konferenzmaterialien
Global

Exploring the stability of feature selection methods across a palette of gene expression datasets

Résumé

Einzelheiten

Aktionen

PDF