Quick answer: A FASTA file is a plain-text file that stores biological sequences: DNA, RNA, or protein. Each record starts with a header line beginning with >, followed by the sequence itself written as single letters, one line or many. It is the single most common file format in bioinformatics, and every sequence tool you will ever use can read it.

What does a FASTA file actually look like?

Here is a complete, valid FASTA file with two records:

>seq1 Human insulin gene fragment
ATGGCCCTGTGGATGCGCCTCCTGCCCCTGCTGGCGCTGCTGGCCCTCTGG
GGACCTGACCCAGCCGCAGCCTTTGTGAACCAACACCTGTGCGGCTCACAC
>seq2 Saccharomyces cerevisiae GAL4 fragment
CTTACCTTCATCGATGACAACAATCAGAACCTTCCCTTCCTGCTGACTTC
ATCAAGCATTTCATCAGAAACTTGAACGTGGAACGCTTGGTGAAGCATTG

The anatomy is simple. The header line starts with > and holds a sequence identifier plus any description you like. Every line after it, until the next > or the end of the file, is the sequence: one letter per nucleotide (A, T, G, C) or one per amino acid for proteins. Blank lines between records are allowed and ignored.

DNA FASTA vs protein FASTA

Both use identical structure; the alphabet tells you which you have. DNA sequences use A, T, G, C (plus N for unknown bases). Protein sequences use the one-letter amino acid codes, which include most letters of the alphabet, so you will see characters like K, Q, W and L that never appear in DNA. Ambiguity codes exist too: R means purine (A or G), Y means pyrimidine (C or T), and so on.

Reading a FASTA file with Python and Biopython

The fastest way to work with FASTA files is Biopython. Install it with pip install biopython, then:

from Bio import SeqIO

for record in SeqIO.parse("sequences.fasta", "fasta"):
    print(record.id, len(record.seq))

That tiny pattern is the foundation of nearly every sequence-analysis script: loop over records, read the identifier and the sequence, and do your analysis per record. Our Python for bioinformatics guide builds on exactly this starting point, and our FASTA and FASTQ formats guide covers how the two formats relate.

Where FASTA files come from

  • Databases: download reference sequences from NCBI, Ensembl or UniProt in FASTA format with one click.
  • Sequencing pipelines: raw instrument output arrives as FASTQ; quality trimming and filtering steps commonly produce cleaned FASTA files.
  • Your own work: any script that writes sequences can emit FASTA, since it is just text.

One caution: FASTA stores only the sequence and a free-text header. It carries no quality scores. When per-base quality matters, such as right after sequencing, use FASTQ, which wraps each base with a quality character. When you just need the sequence itself for alignment, BLAST searches or translation, FASTA is the format the whole field shares.

Frequently asked questions

What does FASTA stand for?

FASTA is named after the FAST-All algorithm, an early and influential sequence-search program from 1988 by David Lipman and William Pearson. The format was created to exchange sequences for that tool and outlived it.

What is the difference between FASTA and FASTQ?

FASTQ stores each sequence together with per-base quality scores, one quality character per base. FASTA stores only the sequence. Raw sequencer output is FASTQ; cleaned or reference sequences are usually FASTA.

Can a FASTA file hold multiple sequences?

Yes. A multi-FASTA file simply contains several records, each starting with its own header line beginning with the greater-than symbol. This is how BLAST databases and reference genomes are distributed.

How do I know if my FASTA file is DNA or protein?

Look at the letters. If you only see A, T, G, C and N it is DNA. If you see letters like K, Q, W, L, E or F it is protein, since protein sequences use the full one-letter amino acid alphabet.

Is there a line length limit in FASTA files?

Traditionally lines were wrapped at 60 or 80 characters, and many tools still expect that. Modern parsers accept single-line sequences, but wrapping keeps compatibility with older software.