WIPIVERSE

Cancer Genome Project

The Cancer Genome Project (CGP) is a large‑scale research initiative that aims to identify the genetic mutations driving the development and progression of human cancers. Established in 2005 at the Wellcome Trust Sanger Institute (now the Wellcome Sanger Institute) in the United Kingdom, the project became a foundational component of the International Cancer Genome Consortium (ICGC), which coordinates worldwide efforts to catalogue somatic mutations across many tumor types.

Objectives

  • Systematically sequence the genomes of cancer specimens and matched normal tissue to detect somatic alterations.
  • Catalogue the spectrum of point mutations, insertions/deletions, copy‑number changes, structural rearrangements, and epigenetic modifications.
  • Distinguish driver mutations that confer selective growth advantages from passenger mutations that are biologically neutral.
  • Provide an open‑access database of validated cancer‑associated genomic alterations for use by researchers and clinicians.

Key Milestones

  • 2005–2009: Initial pilot studies sequenced the genomes of 24 breast and colorectal cancers, demonstrating that whole‑genome sequencing could reveal previously uncharacterized driver genes.
  • 2010: Publication of the first comprehensive catalogue of somatic mutations in breast cancer, identifying recurrent mutations in genes such as PIK3CA and GATA3.
  • 2012: Expansion to cover additional tumor types, including lung, ovarian, and pancreatic cancers, with coordinated projects across multiple international sites.
  • 2014–2020: Integration of whole‑exome sequencing, transcriptome profiling, and DNA methylation data, leading to the identification of novel therapeutic targets such as BRCA1/2 alterations in ovarian cancer and IDH1/2 mutations in glioma.
  • 2021 onward: Ongoing contribution to the ICGC’s “Pan‑Cancer” analyses, providing data for meta‑analyses that compare mutation patterns across more than 50 cancer types.

Methodology

  1. Sample Collection: Tumor specimens are obtained from hospitals and biobanks with patient consent and ethical approval. Matched normal tissue (often blood) provides a reference for germline DNA.
  2. Sequencing Platforms: High‑throughput next‑generation sequencing (Illumina, and later also Pacific Biosciences and Oxford Nanopore) is employed for whole‑genome, whole‑exome, and targeted panels.
  3. Bioinformatic Pipeline: Raw reads undergo quality control, alignment to the human reference genome (GRCh38), and variant calling using validated algorithms (e.g., MuTect2, Strelka). Structural variants are detected with tools such as Manta or DELLY.
  4. Annotation & Validation: Variants are annotated against databases (COSMIC, ClinVar) and filtered for functional impact. Recurrent driver candidates are experimentally validated in cell‑line and animal models.
  5. Data Release: Processed data and accompanying metadata are deposited in public repositories (ICGC Data Portal, European Genome‑phenome Archive) under controlled access where required.

Impact and Applications

  • Clinical Translation: The CGP’s catalogs have informed precision‑medicine trials, enabling genotype‑driven patient stratification for targeted therapies (e.g., PARP inhibitors for BRCA1/2‑mutated tumors).
  • Biomarker Development: Recurrent mutation signatures derived from CGP data have been used to infer exposure to carcinogens (e.g., tobacco‑related signatures in lung cancer).
  • Drug Discovery: Identification of previously uncharacterized oncogenes and tumor suppressors has guided preclinical drug‑screening programs.
  • Educational Resources: The CGP provides training modules and analytic pipelines for graduate students and early‑career researchers in cancer genomics.

Collaboration and Funding
The CGP is funded primarily by the Wellcome Trust, the UK Medical Research Council, Cancer Research UK, and partner institutions within the ICGC. Collaborative links include the Cancer Genome Atlas (TCGA) in the United States, the European Pan‑Cancer Project, and national cancer genomics programs in Japan, Australia, and other countries.

Data Accessibility
Researchers can request access to raw sequencing data and processed mutation calls through the ICGR (International Cancer Genome Repository) portal. Summary statistics, including mutation frequency tables and annotated driver lists, are publicly downloadable without restriction.

Current Status
As of 2024, the Cancer Genome Project continues to expand its repertoire of tumor types and methodological depth, incorporating single‑cell sequencing, spatial transcriptomics, and long‑read technologies to capture intra‑tumoral heterogeneity and complex genomic rearrangements. The ongoing integration of these datasets aims to refine the understanding of cancer evolution and support the next generation of precision oncology.

Browse

More topics to explore

    Browse all articles