Quick answer: Multiple sequence alignment (MSA) lines up three or more DNA or protein sequences so that homologous positions are stacked in columns. It reveals conserved regions, evolutionary relationships and functional domains, and is usually performed with tools such as Clustal Omega, MUSCLE or MAFFT.

In bioinformatics, comparing sequences is one of the most common tasks. When you compare two sequences it is called pairwise alignment; when you compare three or more, it becomes a multiple sequence alignment. This easy guide explains what MSA is, why it matters, and which beginner-friendly tools you can use.

What is a multiple sequence alignment?

A multiple sequence alignment takes three or more sequences and arranges them so that evolutionarily related residues occupy the same column. Gaps are introduced to shift columns, and letters that line up vertically are assumed to share a common ancestor at that position.

Here is a tiny example of three protein sequences:

seq1   M A P K L Y D Q E
seq2   M A P - L Y D Q E
seq3   M A P K L Y D - E

The dashes are gaps. Because the -(gap) characters line up in the same columns as K and Q in other sequences, you can see at a glance where insertions or deletions (indels) happened during evolution.

Why perform a multiple sequence alignment?

  • Conserved regions: columns where most sequences share the same residue highlight functionally important parts of a protein.
  • Phylogenetic analysis: an MSA is the required input for building evolutionary trees.
  • Homology searching: aligned families make it easier to spot distant relatives.
  • Structure prediction: known structures of some members help predict the structure of others.

Pairwise vs multiple sequence alignment

Pairwise alignment compares two sequences, for example when you run a single query through BLAST. Multiple sequence alignment compares at least three sequences and is more computationally expensive. Pairwise methods are exact (Smith-Waterman, Needleman-Wunsch), while multiple alignments rely on progressive or iterative heuristics to stay fast.

How do MSA tools like Clustal Omega work?

Most tools use a technique called progressive alignment:

  1. Every pair of sequences is aligned to build a similarity matrix.
  2. The matrix is converted into a guide tree, a rough evolutionary tree.
  3. Sequences are aligned in the order dictated by the tree, adding one sequence or aligned group at a time.
  4. A final step refines the result to reduce errors.

Clustal Omega is the current version of the classic Clustal family and is fast even for thousands of sequences. MUSCLE and MAFFT are fast alternatives with slightly different alignment strategies.

Scoring and substitution matrices

To decide whether two residues should be aligned, tools use a substitution matrix. For proteins, the commonly used matrices are BLOSUM (especially BLOSUM62) and PAM. The matrix gives a score for aligning every pair of amino acids; conserved pairs get high scores, rare substitutions get low or negative scores. Combined with a gap penalty, the tool can compute the highest-scoring alignment.

Beginner-friendly ways to run an MSA

  • EMBL-EBI Clustal Omega web server: paste your FASTA sequences, click submit, and download the alignment. This is the easiest entry point.
  • UniProt aligned region view: if your protein is in a database, precomputed alignments are often one click away.
  • Seaview or Jalview: desktop viewers let you open, color and edit alignments.

For a quick refresher on what sequences look like before you align them, see our easy read on biological sequences.

Frequently asked questions

What is the difference between pairwise and multiple sequence alignment?

Pairwise alignment compares exactly two sequences, while multiple sequence alignment compares three or more at once to expose common columns and conserved regions.

Which MSA tool should a beginner use first?

Clustal Omega on the EMBL-EBI web server is the best starting point because it needs no installation and gives clean output fast.

What do the dashes in an alignment mean?

Dashes are gaps that represent insertions or deletions relative to the other sequences, and they let homologous residues line up in the same columns.

What is BLOSUM62?

BLOSUM62 is a substitution matrix that scores how likely each pair of amino acids is to replace each other, and it is the default in many tools.

Can I build a phylogenetic tree from an MSA?

Yes. A multiple sequence alignment is exactly the input needed to build a phylogenetic tree of the sequences.