WIPIVERSE

GENCODE

GENCODE is an international collaborative project that aims to provide high‑quality, comprehensive annotation of gene features in the human (Homo sapiens) and mouse (Mus musculus) genomes. The project is coordinated by the GENCODE consortium, which includes researchers from the Wellcome Sanger Institute, the European Bioinformatics Institute (EMBL‑EBI), the National Center for Biotechnology Information (NCBI), and other academic and research institutions.

Purpose and Scope
The primary objective of GENCODE is to identify and annotate all protein‑coding genes, non‑coding RNA genes, and pseudogenes, as well as to delineate splice variants, transcription start sites, and other functional elements. The annotations are intended to serve as a reference for genomic studies, including functional genomics, transcriptomics, and clinical research.

Methodology
GENCODE integrates multiple lines of evidence to produce its gene models:

  1. Experimental data – including RNA‑seq, cap analysis of gene expression (CAGE), and long‑read sequencing.
  2. Computational prediction – using ab initio gene‑finding algorithms and comparative genomics.
  3. Manual curation – expert annotators review and refine automated predictions, resolve ambiguities, and incorporate literature‑derived evidence.

Data Releases and Accessibility
GENCODE releases are periodically updated (e.g., GENCODE v45 for human, vM28 for mouse). Each release is assigned a version number and includes comprehensive annotation files in standard formats such as GTF (Gene Transfer Format) and GFF3. The data are freely available through the GENCODE website, Ensembl genome browser, UCSC Genome Browser, and via FTP servers.

Relationship to Other Resources
GENCODE annotations are incorporated into the Ensembl genome annotation pipeline and serve as the primary transcript set for the Ensembl/GENCODE gene annotation track. They are also used by other databases, such as RefSeq and the UCSC Known Genes track, as a benchmark for gene annotation quality.

Impact and Applications
The GENCODE dataset underpins a wide range of biomedical research, including:

  • Identification of disease‑associated variants in genome‑wide association studies (GWAS).
  • Interpretation of RNA‑seq data for differential expression and isoform analysis.
  • Development of diagnostic panels and therapeutic targets in precision medicine.
  • Comparative genomics studies across vertebrate species.

Funding and Governance
GENCODE is funded by a combination of governmental research agencies (e.g., the Wellcome Trust, the European Union’s Horizon programmes) and institutional support. Governance is provided by a steering committee that oversees project direction, data release policies, and community engagement.

References

  • Harrow, J. et al. (2012). GENCODE: the reference human genome annotation for the ENCODE project. Genome Research, 22(9), 1760‑1774.
  • Frankish, A. et al. (2019). GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Research, 47(D1), D766‑D773.

(All information reflects publicly available sources up to the knowledge cutoff date of September 2021.)

Browse

More topics to explore

    Browse all articles