| EMBL European Nucleotide Archive (ENA) | The European Nucleotide Archive (ENA) provides a comprehensive record of the world’s nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation. | https://www.ebi.ac.uk/ena/browser/ |
| NCBI Sequence Read Archive (SRA) | The SRA is NIH's archive of high-throughput sequencing data and is part of the International Nucleotide Sequence Database Collaboration (INSDC) that includes the NCBI Sequence Read Archive (SRA), the European Bioinformatics Institute (EBI), and the DNA Database of Japan (DDBJ). Data submitted to any of the three organizations are shared among them. SRA is a primary repository for functional genomics data such as RNA-Seq, ChIP-Seq, and Methyl-Seq. | https://www.ncbi.nlm.nih.gov/sra/ |
| Gene Expression Omnibus (GEO) | GEO is a major public functional genomics data repository that archives high-throughput microarray and next-generation sequence functional genomic datasets. GEO stores the raw data, processed files, and metadata | https://www.ncbi.nlm.nih.gov/geo/ |
| Biostudies | Biostudies is a database dedicated to storing and organizing multimodal data from biological studies and relevant links to other databases. Biostudies also hosts functional genomics data from ArrayExpress, including RNA and protein expression and methylation profiling data. A central concept is a study, which contains metadata, including annotations, protocols, processed data and raw data. | https://www.ebi.ac.uk/biostudies/ |
| The Encyclopedia of DNA Elements (ENCODE) | ENCODE is a public resource containing comprehensive data stemming from efforts to map all functional DNA elements present in the human genome. ENCODE integrates and annotates data on genes, transcripts, and transcriptional regulatory regions, as well as their chromatin states and DNA methylation patterns. | https://www.encodeproject.org/ |
| DNA Data Bank of Japan (DDBJ) | DDBJ collects nucleotide sequence data along with their study and sample information as a member of INSDC (International Nucleotide Sequence Database Collaboration). | https://www.ddbj.nig.ac.jp/ |
| Reference Sequence (RefSeq) | The RefSeq database collects a comprehensive, integrated, non-redundant, set of sequences, including genomic DNA, transcripts, and proteins. RefSeq records are subjected to an automated processing pipeline combined with manual curation component to provide a robust set of reference sequences of genomic, transcript and protein data across a wide range of taxa. | https://www.ncbi.nlm.nih.gov/refseq/ |
| Genome Sequence Archive (GSA) | GSA is a repository for archiving raw sequence reads in the National Genomics Data Center, China National Center for Bioinformation (CNCB-NGDC). | https://ngdc.cncb.ac.cn/gsa/ |
| Ensembl | Ensembl is a widely used open resource that integrates and analyses publicly available genomics data from both eukaryotes and prokaryotes. It provides genome assemblies and rich annotations—such as genes, transcripts, proteins, regulatory elements, genetic variants, and comparative genomics (e.g., homology and gene trees)—that can be explored through an interactive genome browser and downloaded in standard formats. Ensembl is maintained by EMBL European Bioinformatics Institute. | https://www.ensembl.org/info/about/index.html |
| GenBank | GenBank is an open access database of all publicly available nucleotide sequences and their protein translations and annotations | https://www.ncbi.nlm.nih.gov/genbank/ |
| UniProt | UniProtKB is one of the most widely used protein databases. It consists of an expertly curated component and a non-peer-reviewed component. It contains hundreds of thousands of protein descriptions, including function, domain structure, subcellular location, post-translational modifications, and functionally characterized variants. | http://www.ebi.ac.uk/uniprot/ |
| The European Genome-phenome Archive (EGA) | The European Genome-phenome Archive (EGA) is a service for permanent archiving and sharing of personally identifiable genetic, phenotypic, and clinical data generated for biomedical research projects or in research-focused healthcare systems. | https://ega-archive.org/ |
| The Cancer Genome Atlas (TCGA) | TCGA is a collaborative cancer genomics program led by the US National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI) that systematically characterized the molecular features of many human cancers. TCGA generated multi-omics datasets—such as DNA sequencing for mutations, copy-number changes, gene expression, DNA methylation, and other molecular and clinical annotations. It has generated comprehensive, multi-dimensional maps of tumor tissue and matched normal tissues across 33 types of cancers. | http://cancergenome.nih.gov/ |
| Genotype-Tissue Expression (GTEx) | GTEx is a comprehensive resource and tissue biobank that aims to map the genetic effects on the transcriptome across diverse human tissues and to link these regulatory mechanisms to trait and disease. GTEx provides molecular assays, including whole-genome sequencing, whole-exome sequencing, and RNA-seq. | https://gtexportal.org/home/ |
| RCSB Protein Data Bank (PDB) | PDB is a global repository of the experimentally determined 3D structures of large biological molecules, including proteins, DNA and RNA. It is managed by the Worldwide Protein Data Bank organization (wwPDB; wwpdb.org). | https://www.rcsb.org/ |
| UCSC Genome Browser | The UCSC Genome Browser is an interactive website providing access to genome sequence data from a wide range of vertebrate and invertebrate species and major model organisms, integrated with a large collection of aligned annotations. | https://genome.ucsc.edu/ |