Models for genomic data
Elementary input and output models for genomic data types
Partners involved: Stegle, Hammer, Gagneur
Genomic data sets based on sequencing methods are inherently susceptible to technical variation and interfering factors affecting samples. This applies to input data for machine learning methods, genome sequences and targets such as integer-based estimates of gene expression or fraction-based phenotypes such as splicing and methylation. Although classical statistical methods in genomics exclude these technical factors of variance, most existing deep learning methods ignore this variability and often assume normally distributed errors or ignore technical variation. Based on principles from statistical genomics (DESeq, scaling methods for single cell RNA sequencing data sets), we will develop cost functions that take such variation into account.
