Nada Lavrač,Blaž Škrlj,Marko Robnik-Šikonja
Nada Lavrač
Data preprocessing is an important component of machine learning pipelines, which requires ample time and resources. An integral part of preprocessing is data transformation into the format required by a given learning algorithm. This paper...
An evaluation of machine-learning for predicting phenotype: studies in yeast, rice, and wheat [0.03%]
机器学习预测表型的评估研究——以酵母、水稻和小麦为例
Nastasiya F Grinberg,Oghenejokpeme I Orhobor,Ross D King
Nastasiya F Grinberg
In phenotype prediction the physical characteristics of an organism are predicted from knowledge of its genotype and environment. Such studies, often called genome-wide association studies, are of the highest societal importance, as they ar...
An evaluation of linear and non-linear models of expressive dynamics in classical piano and symphonic music [0.03%]
古典钢琴和交响乐中表情力度的线性与非线性模型的评价
Carlos Eduardo Cancino-Chacón,Thassilo Gadermaier,Gerhard Widmer et al.
Carlos Eduardo Cancino-Chacón et al.
Expressive interpretation forms an important but complex aspect of music, particularly in Western classical music. Modeling the relation between musical expression and structural aspects of the score being performed is an ongoing line of re...
Meta-QSAR: a large-scale application of meta-learning to drug design and discovery [0.03%]
元QSAR:药物设计和发现的元学习大规模应用
Ivan Olier,Noureddin Sadawi,G Richard Bickerton et al.
Ivan Olier et al.
We investigate the learning of quantitative structure activity relationships (QSARs) as a case-study of meta-learning. This application area is of the highest societal importance, as it is a key step in the development of new medicines. The...
Konstantinos Sechidis,Gavin Brown
Konstantinos Sechidis
What is the simplest thing you can do to solve a problem? In the context of semi-supervised feature selection, we tackle exactly this-how much we can gain from two simple classifier-independent strategies. If we have some binary labelled da...
Ioannis Tsamardinos,Giorgos Borboudakis,Pavlos Katsogridakis et al.
Ioannis Tsamardinos et al.
We present the Parallel, Forward-Backward with Pruning (PFBP) algorithm for feature selection (FS) for Big Data of high dimensionality. PFBP partitions the data matrix both in terms of rows as well as columns. By employing the concepts of p...
NhatHai Phan,Xintao Wu,Dejing Dou
NhatHai Phan
The remarkable development of deep learning in medicine and healthcare domain presents obvious privacy issues, when deep neural networks are built on users' personal and highly sensitive data, e.g., clinical records, user profiles, biomedic...
Bootstrapping the out-of-sample predictions for efficient and accurate cross-validation [0.03%]
自助法预测出样本数据以实现高效准确的交叉验证
Ioannis Tsamardinos,Elissavet Greasidou,Giorgos Borboudakis
Ioannis Tsamardinos
Cross-Validation (CV), and out-of-sample performance-estimation protocols in general, are often employed both for (a) selecting the optimal combination of algorithms and values of hyper-parameters (called a configuration) for producing the ...
W John Wilbur,Lana Yeganova,Won Kim
W John Wilbur
Schapire and Singer's improved version of AdaBoost for handling weak hypotheses with confidence rated predictions represents an important advance in the theory and practice of boosting. Its success results from a more efficient use of infor...
Amol Pande,Liang Li,Jeevanantham Rajeswaran et al.
Amol Pande et al.
Machine learning methods provide a powerful approach for analyzing longitudinal data in which repeated measurements are observed for a subject over time. We boost multivariate trees to fit a novel flexible semi-nonparametric marginal model ...