Friday, May 11, 2012

Updating the genome: correcting the assembly of 10q11.22


Human GRCh37 patch release 8 contains an update to previously released fix patch HG1211_PATCH.

This encompasses a 3Mb region in GRCh37 between chr10: 46,256,855-49,299,273.

The tile path in the 10q11.22 region has been extensively altered from its previously fragmented state to one where a single gap remains, between BX649215.1 and AC245041.3. The reworking of the tile path in the region has been carried out using clones in the existing build and additional finished clones not previously in GRCh37.

Working with optical map data provided by the Schwartz Lab we have been able to identify errors in the GRCh37 assembly and have consequentially worked to correct them. The optical map analysis also highlighted redundancy in the assembly causing artificial duplication, which has now been addressed within this patch.
Above: Optical map consensus alignments to GRCh37 10q11.22.
Below: Optical map consensus alignments to the fix patch (JH591181.2)
Legend: Pink track: Clone path; Green: Contig gap; Blue: In silico SwaI fragments.
For the aligned optical map consensus Gold: Concordant fragment ; Red: Missing fragment (seen where OM consensus span gap); Grey: Unaligned fragment







The optical map information was consistent with a path problem in this region. The map data suggested that several clones in the region were misplaced and did not represent a valid chromosome structure in this region. In addition to rearranging several clones (including changing the orientation of some clones in the path), 3 finished clones were added to the path and several redundant clones were removed. The new path contains a single gap that we estimate, based on optical mapping, to be about 90 Kb. The figure below shows an alignment of the patch sequence to the current chr10 assembly.
The panel to the left shows an overview of chr. 10. The orange dots represent fix patches we've released and the blue dots represent novel patches. The arrow shows the location of the 10q21 fix patch. To the right, the top panel shows the chr. 10 tiling path (in grey), the annotated RefSeq genes are below that (in green) and the alignment to the fix patch below that (in purple). The bottom panel shows the patch tiling path and alignment to the chromosome. 





Friday, May 4, 2012

Filling in the gaps to better understand human biology


Duplicated segments pose serious problems for the assembly and annotation of the human genome. In the human reference genome there are still large gaps that require specialized efforts to fill. Many of these gaps lie within highly duplicated segments in which the degree of sequence variation among duplicated loci approaches levels of allelic variation. Many people assume that much of the sequence that is still missing from the reference assembly is not very biologically interesting. However, it has become increasingly apparent that the segmental duplications themselves provide the molecular basis for many human genetic disorders. The resolution of these regions is therefore essential for a complete understanding of the genetic basis of human disease. 

Three patches released in GRCh37.p8, that add almost 400Kb of novel sequence, prove the concept that sequence missing so far from the reference genome can be of crucial importance.  The biological story surrounding these sequences can be found in a recent publication from the Eichler lab (Dennis et al., 2012) but here we'll tell you a little bit about how we worked with the Eichler lab to create these assembly patches. 

Figure 1: Ancestral copy of SRGAP2 in 
chimpanzee (left) and human (right). The other 
red ticks on the human chromosome show
the human specific duplications added by 
this effort.

One of the impediments in resolving the complexity of these regions is the diploid nature of the human genome. We recently took advantage of a haploid BAC library resource (CHORI-17) from hydatidiform mole DNA to close gaps and resolve the genomic structure of segmental duplications encompassing highly identical paralogs of SRGAP2, a gene important in cortex development. Hydatidiform moles are conception abnormalities that most often arise from the fertilization of an enucleated ovum by a single X-bearing sperm. Subsequent diploidization results in a 46 XX karyotype in which all allelic variation has been eliminated allowing the unambiguous delineation of duplicated DNA as well as haplotype characterization. Our SRGAP2 sequencing efforts resolved the sequence and structure of 4 copies of the gene on human chromosome 1, three of which represent human-specific duplicate truncations of the original ancestral gene. 


Overall, we added >380 kbp of new sequence previously absent from the human reference genome, including 40 kbp within the conserved ancestral copy of the gene. Additionally, we discovered ~560 kbp of sequence mapped incorrectly either in orientation or position. This region in GRCh37 contained 15 gaps, and now in the new sequence patch, only two gaps remain.  Combined, we generated or corrected 0.4% of human chromosome 1 euchromatic sequence. The sequencing of these genes have made it possible to explore the function of the human-specific duplicate copies, particularly their role in  neurological traits and disorders unique to humans.
SRGAP2A (1q32 region): JH636054.1
SRGAP2B,D (1q21 region): JH636052.1
SRGAP2C (1p12 region): JH636053.1

Figure 2: View of SRGAP2 gene family on chromosome 1 ideogram (with 1q on the left). The arrows show the order and direction of duplication with the estimated time (in millions of years ago) below that. (Dennis et al., 2012)




Friday, April 27, 2012

Updating the Human Reference Assembly, part 1

Talking about updating the human reference assembly, currently GRCh37 (hg19), can elicit groans and howls of protests from genome scientists who have put considerable effort into analyzing a given data set against the reference assembly. To address this concern, we introduced the notion of a 'Genome Patch'; that is, scaffold size sequences that either add additional sequence representation (NOVEL patches) or fix existing problems in the current reference assembly (FIX patches). In this way, we can make our best representation of the assembly available without disrupting the reference chromosome coordinates. We are at our eighth patch release (GRCh37.p8) and we now have 69 FIX patches and 71 NOVEL patches. 


It is the FIX patches we'd like to consider right now. While the patch scaffolds are easy enough to use if you are interested in a single region, most analysis pipelines have not incorporated these sequences and the improved data remains largely unused for whole genome or exome analysis. It is worth noting that NCBI and Ensembl provide gene annotation on many of the patch releases. Doing a major update to the reference assembly (making GRCh38) will allow us to incorporate these FIX patches into the chromosome assembly, making them directly accessible to analysis pipelines.

There are 66 regions (>40.5 Mb) on GRCh37 that are associated with these 69 FIX patches. In addition to other sequence changes that improve the reference assembly, the sequences in these FIX patches provide more than 4.7 Mb of novel sequence.  Adding large amounts of novel sequence, like the 2.6 Mb added by JH636052.1 (1q21 region) is impressive, however, novel sequence is not the only metric to consider when evaluating FIX patches. For example, GL383543.1/NW_003315932.1 (described in HG-544) is a FIX patch for the FAM23A_MRC1 region on NC_000010.10 (chr10: 17613209-18252930) and adds no novel sequence to the reference assembly. Instead, it removes roughly 200 Kb of artificially redundant sequence and closes a gap in the assembly. The alignment of the patch to the chromosome is shown in Figure 1 (below). 

Alignment of FIX patch to chromosome for FAM23A_MRC1 region of chr10.




Figure 1: FAM23A_MRC1 region on chromosome 10: The top panel shows chr10 in GRCh37. The blue/black line at the top represents the sequence, the track below that is the GenBank components used to assemble the chromosome, below that are the NCBI genes, then the alignment of the chromosome to the patch sequence and finally the segmental duplication track. The second panel shows the FIX patch sequence, which has no gap, the genes annotated on the patch and the alignment to the chromosome. The patch removes roughly 200Kb of artificially redundant sequence (meaning the data in the segmental duplication track is an artifact) and corrects the gene annotation in the region, removing two gene models that represent false gene duplications and don't exist in the population. (see full size photo)



The artificial duplication in the assembly not only affects the gene annotation but also has a significant affect on the alignment of short reads as shown in Figure 2 (below), or you can see the alignments

Figure 2: Alignments of 1000 Genomes data the FAM23_MRC1 region on chromosome 10: The 1000 Genome data alignments are in the tracks below the orange bar noting where the artificial duplication exists in the reference assembly. Two low-coverage samples (NA19625 and NA19701) are aligned using BWA and Mosaic respectively. In the two Mosaic tracks, there is a visible drop off in alignment depth. This is less pronounced in the BWA alignments. The red coloring indicates mismatches in the alignments. (see full size photo)


While it is well-recognized that sequences from the reference, particularly missing paralogs (see Sudmant et al., 2010 for more information), have affects on next generation sequence analysis, it should be noted that artificial duplication within the assembly, such as the example shown here, can also significantly impact such analyses. With the latest patch release we have updated 5 such regions (covering 2 Mb). Several other regions that are phenotypically important, but were represented by a mixed haplotype in GRCh37, have also been updated, including the Williams region on chr7 and the 1q21 region on chr1.

We'll talk about other things FIX patches get you in a later post. Additionally, we'll be highlighting some biologically interesting regions!

Tuesday, April 17, 2012

GRCh37.p8 is now available!

The latest patch release for human (patch 8) is now available! For GRCh37.p8, we've released 9 new FIX patches, 1 new NOVEL patch and we updated one FIX patch from a previous release. We'll provide some more information about specific patches in future blog posts, but if you wanted to get the latest data now go to our FTP site.

Tuesday, July 5, 2011

Genome Update: Representing variation in the LRC on chr. 19q13.4


Human GRCh37 patch release 5 includes eight NOVEL patches representing different haplotypes in  the Leukocyte Receptor Complex (LRC) region on chromosome 19q13.4 (GL949746.1, GL949747.1, GL949748.1, GL949749.1, GL949750.1, GL949751.1, GL949752.1, GL949753.1). This region contains multiple clusters of genes belonging to the immunoglobulin superfamily, including killer immunoglobulin-like receptors (KIRs), leukocyte immunoglobulin-like receptors (LILRs) and leucocyte-associated immunoglobulin-like receptors (LAIRs). The LRC complex is of major importance in human disease across a wide context. Research efforts have focused in particular on the KIR cluster, since this ~150kb  region displays extensive haplotypic variation due to both differences in coding sequences and the presence or absence of particular loci. 

Several reports indicated problems with the representation of the LRC region in both NCBI36 and GRCh37. In GRCh37, one improvement was made when the NCBI36 chr. 19 unlocalized scaffold NT_113949.1, which contained a second representation of this region, was determined to be mis-assembled and was excluded from the assembly (tracked in HG-196). However, in both assembly versions, the chromosome 19 sequence for this variable region is derived from multiple clone libraries, suggesting a haplotype representation problem. On-going GRC efforts to replace this region of chromosome 19 in future assembly versions with a new single haplotype  from the CHORI-17 hydatidiform mole library are being tracked in HG-1079. The NOVEL LRC patches that have now been released provide partial representations of the LRC region for eight different haplotypes.

Four of the NOVEL patch LRC haplotypes are derived from the same PGF and COX cell lines that were used in the Major Histocompatibility Complex (MHC) project (7 haplotypes from the MHC project have already been incorporated into GRCh37 as alternate loci: GL000250.1, GL000251.1, GL000252.1, GL000253.1, GL000254.1, GL000255.1, GL000256.1). However, whilst PGF and COX are homozygous for the HLA region of the MHC, they are heterozygous for the KIR region of the MHC, and hence are represented here as PGF1 and 2, and COX1 and 2 (PMID:17092261). The other four LRC haplotypes, named s, t, j and i, are derived from a study by Traherne et al. that identified rare contracted KIR haplotypes in families of European origin (PMID: 19959527).

The sequence coverage of the s, t, j and i haplotypes is limited to the KIR region, whilst that of COX1/2 and PGF 1/2 extends in to the LILR and LAIR clusters. Corresponding manual gene annotation for each of these haplotypes has been generated as part of the Vega project.

Figure 1 (below): Alignment of the 8 LRC region NOVEL patches to GRCh37 chr. 19. The blue bars at top represent the tiling path of chr. 19 (NC_000019.9). Genes annotated on this sequence are shown in green. The gray tracks below represent the alignments: the thin horizontal lines indicate gaps, while the small vertical red bars indicate mismatches. 

Wednesday, March 23, 2011

Updating the genome: the CCL3L1 region of chr17q21

The CCL3L1 and CCL4L1 genes are found in a region of Human chromosome 17q12. These genes encode cytokines and the number of gene copies varies between individuals, with 0-4 copies in European individuals and 3-10 copies in African individuals. Copy number variations of these genes have been associated with various autoimmune diseases, possibly playing a role in rheumatoid arthritis susceptibility [PMID:17604289]. There are conflicting reports concerning how this region influences HIV infection and progression [PMID:15637236 and PMID:19812560].

In the NCBI36 reference, this region was comprised of clones from different libraries, and thus different haplotypes. A user reported that it seemed likely that the selected tiling path, that also contained a gap, did not represent a valid structure at this biomedically important locus. We have tracked the work on this region in HG-75. Despite our best efforts, we could not resolve all problems in this region in time for the release of GRCh37.

Because of the complexity of this region, we chose to produce a new tiling path using a BAC library that has been constructed from a hydatidiform mole library (CHORI-17). Complete hydatidiform moles are the result of a single sperm fertilizing an enucleated egg. The sperm reduplicates to generate two sets of the paternal chromosomes and thus contains DNA from a single haplotype. Clones from this resource have proven very useful for resolving highly duplicated genomic regions. The joins between the newly sequenced mole clones are of excellent quality, so we have a higher degree of confidence in the assembly of the new components than in the old, mixed-haplotype assembly.

 The resultant pathway closes the gap and contains a single copy of each CCL3L1 and CCL4L1, providing a valid allele at this locus. The new pathway also contains five full copies of TBC1D3, two of which flank the CCL genes, and could provide a reasonable explanation for the generation of the null state resulting from the recombination between the CCL flanking copies of TBC1D3. Because of the clinical importance of this region, we released this sequence as part of patch release 2 (GL383560.1).


Figure 1: Alignment of GL383560.1 to chr17 sequence. The top track, represented by the blue line, shows the sequence of GL383560.1. Below that is a track of gene features annotated on this sequence (blue represents transcript features and red represents CDS). The track below this, represented by the gray line with the red vertical bars, is the alignment to chr17 (CM000679.1). Vertical red bars show mismatches, the then red lines are gaps in the CM000679.1 sequence. The annotation on chr17 is projected below this alignment so that the resulting change in annotation can be seen.

Wednesday, October 13, 2010

Zebrafish genome joins GRC

The Genome Reference Consortium has expanded to take over the maintenance and improvement of the zebrafish genome sequence.

The zebrafish genome was sequenced by the Wellcome Trust Sanger Institute. The GRC has recognised the high quality of the current zebrafish genome assembly Zv9, and its importance as a model organism genome. The consortium has therefore committed to further improve the zebrafish genome sequence and to maintain it for an unlimited time in the future. Work has started to replace whole genome shotgun sequence with high quality finished clone sequence, close remaining gaps and place yet unlocalized sequence.

The GRC's web pages are available to the community to report genome issues and track progress of the clone path development and assembly generation. 

Please send us your questions and comments to zfish-help@sanger.ac.uk or to the GRC directly. Your input is highly valued and helps us to improve the support for your research!

Photo courtesy of Lukas Roth
Links:
Zebrafish Genome Project home page
Zebrafish assembly Zv9 in pre Ensembl
Zebrafish clone path in Vega