Quick answer: Bioinformatics databases are organized collections of biological data, from DNA and protein sequences to structures and gene annotations. The most important ones for beginners are NCBI GenBank and RefSeq for sequences, UniProt for proteins, and Ensembl for genomes and genes.
When you finish sequencing or when you want to identify a sequence, the next stop is almost always a database. Bioinformatics databases store the world’s DNA, RNA, protein, structure and literature data and make them searchable online. This easy guide walks through the main databases every beginner should know and when to use each one.
What makes a bioinformatics database?
A bioinformatics database is simply a large, organized collection of biological data with search and download tools. Entries are usually annotated with identifiers, references, cross-links to other databases and supporting evidence. Central databases such as NCBI, ENA and DDBJ synchronize the same raw sequences among themselves, so a sequence can exist in more than one place with the same accession number.
NCBI: the sequence powerhouse
The National Center for Biotechnology Information (NCBI) hosts the biggest DNA and RNA sequence archive in the world. Its key collections are:
- GenBank: the large public repository of submitted nucleotide sequences with annotations.
- RefSeq: curated, non-redundant reference sequences.
- BLAST databases: prebuilt collections you search with the BLAST tool described on this site.
- PubMed: the biomedical literature database, covering over 35 million articles.
UniProt: the protein knowledge hub
UniProt is the central resource for protein sequences and function. It combines:
- UniProtKB/Swiss-Prot: manually curated, well-reviewed entries with rich annotations.
- UniProtKB/TrEMBL: automatically annotated entries, large but with less expert review.
- UniRef: clustered sequences to remove redundancy.
- UniParc: an archive of all protein sequences.
If you want to find a protein’s function, domains, subsections and cross-references, UniProt is usually your first stop.
Ensembl: genomes made viewable
Ensembl (for vertebrates) and its sibling Ensembl Genomes (for plants, fungi, bacteria and others) provide annotated genome browsers. You can view a chromosome, zoom into a gene, inspect its transcripts and exons, and download its sequence or annotation files such as GFF and FASTA. This is where beginners typically look up “what genes does my organism have”.
Other databases worth knowing
- Protein Data Bank (PDB): experimental 3D structures of proteins and other molecules.
- InterPro: classifies proteins into families and detects domains.
- KEGG: pathways, enzymes and metabolic networks.
- Gene Ontology (GO): a controlled vocabulary describing gene functions across species.
- ENA and DDBJ: the European and Japanese sequence archives, which mirror GenBank content.
Which database should I use and when?
| Your question | Database |
|---|---|
| Identify an unknown DNA sequence | NCBI GenBank via BLAST |
| Look up a protein’s function | UniProt |
| Browse a genome and its genes | Ensembl |
| Check the 3D structure of a protein | Protein Data Bank |
| Find papers and citations | PubMed |
How to search effectively
- Use the accession number when you know it, for example NM_000518 for the human beta globin mRNA.
- Search by gene symbol or protein name when you do not have a number.
- Refine results with species filters to avoid cross-species matches.
- Download records in FASTA format for use in alignment and file-format practice.
Struggling with careers or whether you need to code? Our posts on bioinformatics job scope and whether bioinformatics needs programming answer those questions.
Frequently asked questions
Which bioinformatics database should a beginner learn first?
NCBI is the best starting point because it hosts DNA sequences, BLAST and PubMed, so most workflows begin and end there.
What is the difference between GenBank and RefSeq?
GenBank stores submitted sequences with the original annotations, while RefSeq is a curated, non-redundant reference set that is more stable.
Are the major sequence databases identical?
GenBank, ENA and DDBJ exchange data daily, so the same raw sequence usually appears with the same accession number in all three.
Do I need an account to search these databases?
No. Searching and downloading from NCBI, UniProt and Ensembl works without an account, though logging in adds features like saved searches.

