Project title:

Modelling Allelic ChIP-seq data using statistical analysis and MMDiff2


Krister Jazz Urog

Defense Year: 2021-2022

Due to the advancement of high-throughput sequencing, genomic data has been increasing in a staggering rate. This fast-paced data accumulation constitutes for the need of tools and algorithms that will be able to process and analyze these datasets efficiently and meaningfully. In particular, ChIP-seq experiments are widely used to analyze protein-DNA interactions that is useful for the epigenetics field. In an allelic ChIP-seq data, one of the major problem is determining which genomic reads comes from the maternal or paternal allele. Allelic calls on a ChIP-seq experiment throw away most of the reads resulting to a very small amount of allelic dataset compared to the actual ChIP-seq data. Thus, modelling these allelic ChIP-seq datasets given a few labeled datasets becomes a challenging task. This is also proven to be a tedious task following a wide set of tools required to for a complete ChIP-seq analysis, the large volume of data that we have to analyze, as well as the scarcity of the labeled dataset.

In summary, we were able to implement a pipeline that incorporates a semi-supervised classification algorithm to impute data from a ChIP-seq experiment. Here we present a novel method to model allelic ChIP-seq data which should be computationally expensive without parallel computing, but done in an optimized way using high performance computing tools and task parallelization. This is done by careful integration of the wide set of tools for high-throughput sequencing such as samtools, bedtools, pysam, and bed-tools. Parallelization is done using dask and a preprocessing step to split the data for embarassingly parallel tasks. MMDiff2 results shows an overwhelmingly large number of differential peaks from the imputation, calling for better biological and statistical representations for the algorithm.

Thesis not available.