Project description:Genome graphs, including the recently released draft human pangenome graph, can represent the breadth of genetic diversity and thus transcend the limits of traditional linear reference genomes. However, there are no genome-graph-compatible tools for analyzing whole genome bisulfite sequencing (WGBS) data. To close this gap, we introduce methylGrapher, a tool tailored for accurate DNA methylation analysis by mapping WGBS data to a genome graph. Notably, methylGrapher can reconstruct methylation patterns along haplotype paths precisely and efficiently. To demonstrate the utility of methylGrapher, we analyzed the WGBS data derived from five individuals whose genomes were included in the first Human Pangenome draft as well as WGBS data from ENCODE (EN-TEx). Along with standard performance benchmarking, we show that methylGrapher fully recapitulates DNA methylation patterns defined by classic linear genome analysis approaches. Importantly, methylGrapher captures a substantial number of CpG sites that are missed by linear methods, and improves overall genome coverage while reducing alignment reference bias. Thus, methylGrapher is a first step towards unlocking the full potential of Human Pangenome graphs in genomic DNA methylation analysis.
Project description:Pangenome arrays contain DNA oligomers targeting several sequenced reference genomes from the same species. In microbiology these can be employed to investigate the often high genetic variability within a species by comparative genome hybridization (CGH). The biological interpretation of pangenome CGH data depends on the ability to compare strains at a functional level, particularly by comparing the presence or absence of orthologous genes. Due to the high genetic variability, available genotype-calling algorithms can not be applied to pangenome CGH data. Therefore, we have developed the algorithm PanCGH that incorporates orthology information about genes to predict the presence or absence of orthologous genes in a query organism using CGH arrays that target the genomes of sequenced representatives of a group of microorganisms. PanCGH was tested and applied in the analysis of genetic diversity among 39 Lactococcus lactis strains from three different subspecies (lactis, cremoris, hordniae) and isolated from two different niches (dairy and plant). Clustering of these strains using the presence/absence data of gene orthologs revealed a clear separation between different subspecies and reflected the niche of the strains. Keywords: CGH, CGH analysis, orthology, Lactococcus lactis
2008-12-31 | GSE12638 | GEO
Project description:HiC data of pepper accessions for graph pangenome of Capsicum genus
Project description:To investigate the functional and adaptive significance of allele-specific candidate cis-regulatory elements (cCREs), we generated ATAC-seq on liver tissue from two 25-day-old F1 hybrid mice, bred from a CC032 mother and a CC072 father. Reads were mapped to the Collaborative Cross pangenome graph and phased into CC032 and CC072 alleles. Because no peak caller currently works directly on minigraph-cactus graphs, reads were surjected to GRCm38 for peak calling with MACS2. Counting phased reads on both alleles at 102,412 reproducible cCREs and testing for differential accessibility identified 1,805 allele-specific cCREs (1.7%) between the maternal and paternal genomes. About 72% of these carry SNVs and/or SVs from one allele, and they are enriched in promoters, introns and intergenic regions as well as in SINE, LINE and LTR transposable-element subfamilies.