Quick answer: FASTA and FASTQ are the two most common text file formats in bioinformatics. A FASTA file stores a sequence with a single header line that starts with >, while a FASTQ file stores four lines per read: a header starting with @, the sequence, a + separator, and a quality-score string of equal length.

If you are just starting in bioinformatics, the very first thing you will touch after a sequencing run is a file format. Before any analysis can begin, you need to open, read and validate the sequence data. This easy guide explains FASTA and FASTQ so you can tell them apart at a glance and handle both correctly.

What is a FASTA file?

FASTA is a simple, human-readable text format that stores one or more DNA, RNA or protein sequences. Each record has two parts:

  • a header line starting with the > character followed by a description, and
  • one or more lines of sequence letters (A, C, G, T for DNA; the 20 amino acid letters for proteins).

A typical FASTA record looks like this:

>seq1 human beta globin
ATGGTGCACCTGACTCCTGAGGAGAAGTCTGCCGTTACTGCCCTGTGGG
GCAAGGTGAACGTGGATGAAGTTGGTGGTGAGGCCCTGGGCAGGTTGGT
>seq2 mouse beta globin
ATGGTGCACCTGACTGATGCTGAGAAGGCTGCTGTCTCTGCCTGTGGG

The identifier is the word right after the > symbol (for example, seq1). Many databases such as NCBI GenBank store their sequence records in a FASTA-like layout because it is compact and easy to parse.

What is a FASTQ file?

FASTQ is the standard output format of modern high-throughput sequencing machines. It stores the sequence and a per-base quality score, which tells you how confident the machine is about each nucleotide call. Every read occupies exactly four lines:

  • line 1: header, starting with @
  • line 2: the raw sequence letters
  • line 3: a separator, usually just +
  • line 4: a quality string with the same length as the sequence
@SRR001666.1 071112_SLXA-EAS1_s_7:5:1:817:345 length=36
GGGTGATGGCCGCTGCCGATGGCGTCAAATCCCAC
+
IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIII

The fourth line looks like random characters, but each symbol maps to a number. That number is the Phred quality score for the base above it, giving you a measure of sequencing error.

How are quality scores written?

FASTQ quality scores use the Phred scale. A Phred score of 30 (written as the character I in the standard Illumina encoding) means the probability that the base call is wrong is about 1 in 1000, so it is 99.9% accurate. Higher is better. The score characters are usually encoded with the ASCII offset 33, which is why FASTQ quality lines look like punctuation and letters.

FASTA vs FASTQ: the quick comparison

FeatureFASTAFASTQ
Header marker>@
Stores qualityNoYes
Used forReference or assembled sequences, proteinsRaw sequencing reads
Lines per record2 or moreExactly 4
Readable by humansVery easyQuality line looks cryptic

You can convert between the two formats, but you lose information: converting FASTQ to FASTA keeps the sequence and drops the quality scores, which is only safe if you no longer need to filter reads by quality.

Where do these files come in a typical workflow?

A typical genomics workflow starts with sequencing, which produces FASTQ files of raw reads. The reads are then cleaned, aligned to a genome, and the results are often stored against a FASTA reference. If you are comparing new sequences against databases, tools such as BLAST accept FASTA input directly.

Common mistakes beginners make

  • Using @ for a FASTA header by mistake, which confuses parsers that expect >.
  • Mixing sequence and quality lines in FASTQ, so the quality string no longer matches the sequence length.
  • Adding trailing spaces or blank lines inside records, which some tools reject.
  • Using protein letters in a DNA-only tool.

If you want a refresher on the molecules these files describe, our easy read on nucleic acids and nucleosides and nucleotides posts are good starting points.

Frequently asked questions

Which is better, FASTA or FASTQ?

Neither is better overall. FASTA is simpler and stores just sequence, while FASTQ adds quality scores, which makes it the right choice for raw sequencing reads.

Can I open FASTA and FASTQ files in Excel?

You can open them as plain text, but the fixed four-line structure of FASTQ and long sequences make a plain text editor a better choice for daily work.

Why does the FASTQ quality line look like symbols?

The quality line is the Phred score shifted by ASCII offset 33, so values show up as punctuation and letter characters instead of plain numbers.

Is FASTA used for protein sequences too?

Yes. FASTA works for both nucleotide and protein sequences, using the 20 standard amino acid letters for proteins.