Restriction Enzyme Recognition Sequences Reference & Cut-Site Finder
Recognition sequences, cut sites, and end types for 30 commonly used restriction enzymes, plus a site-finder: paste a sequence and see where each enzyme cuts and how many fragments it produces.
30 common restriction enzymes
| Enzyme | Recognition site (^ = cut) | Length | Ends | Overhang (nt) | Mean spacing in random DNA |
|---|---|---|---|---|---|
| AgeI | A^CCGGT |
6 | 5′ overhang | 4 | 4,096 bp |
| AluI | AG^CT |
4 | blunt | 0 | 256 bp |
| ApaI | GGGCC^C |
6 | 3′ overhang | 4 | 4,096 bp |
| AvrII | C^CTAGG |
6 | 5′ overhang | 4 | 4,096 bp |
| BamHI | G^GATCC |
6 | 5′ overhang | 4 | 4,096 bp |
| BglII | A^GATCT |
6 | 5′ overhang | 4 | 4,096 bp |
| BsrGI | T^GTACA |
6 | 5′ overhang | 4 | 4,096 bp |
| ClaI | AT^CGAT |
6 | 5′ overhang | 2 | 4,096 bp |
| DpnI | GA^TC |
4 | blunt | 0 | 256 bp |
| EcoRI | G^AATTC |
6 | 5′ overhang | 4 | 4,096 bp |
| EcoRV | GAT^ATC |
6 | blunt | 0 | 4,096 bp |
| HaeIII | GG^CC |
4 | blunt | 0 | 256 bp |
| HindIII | A^AGCTT |
6 | 5′ overhang | 4 | 4,096 bp |
| HpaII | C^CGG |
4 | 5′ overhang | 2 | 256 bp |
| KpnI | GGTAC^C |
6 | 3′ overhang | 4 | 4,096 bp |
| MluI | A^CGCGT |
6 | 5′ overhang | 4 | 4,096 bp |
| MspI | C^CGG |
4 | 5′ overhang | 2 | 256 bp |
| NcoI | C^CATGG |
6 | 5′ overhang | 4 | 4,096 bp |
| NdeI | CA^TATG |
6 | 5′ overhang | 2 | 4,096 bp |
| NotI | GC^GGCCGC |
8 | 5′ overhang | 4 | 65,536 bp |
| PstI | CTGCA^G |
6 | 3′ overhang | 4 | 4,096 bp |
| SacI | GAGCT^C |
6 | 3′ overhang | 4 | 4,096 bp |
| SalI | G^TCGAC |
6 | 5′ overhang | 4 | 4,096 bp |
| SmaI | CCC^GGG |
6 | blunt | 0 | 4,096 bp |
| SpeI | A^CTAGT |
6 | 5′ overhang | 4 | 4,096 bp |
| SphI | GCATG^C |
6 | 3′ overhang | 4 | 4,096 bp |
| TaqI | T^CGA |
4 | 5′ overhang | 2 | 256 bp |
| XbaI | T^CTAGA |
6 | 5′ overhang | 4 | 4,096 bp |
| XhoI | C^TCGAG |
6 | 5′ overhang | 4 | 4,096 bp |
| XmaI | C^CCGGG |
6 | 5′ overhang | 4 | 4,096 bp |
Enzymes sharing a recognition site
| Site | Enzymes | Same cut position | Relationship |
|---|---|---|---|
| CCGG | HpaII, MspI | Yes | Isoschizomers |
| CCCGGG | SmaI, XmaI | No | Neoschizomers |
Where this data comes from
Recognition sequences and cut sites are parsed programmatically from the raw data files of REBASE (The Restriction Enzyme Database, Roberts RJ et al., http://rebase.neb.com, version 609, 2026-08-27). No manual transcription was involved.
The end type, overhang length, and average spacing columns are calculated, not copied: for a recognition sequence of length L with the top-strand cut at position k, overhang length = L − 2k. Positive values indicate a 5′ overhang, negative values a 3′ overhang, and zero means blunt ends. Average spacing is 4^L — the expected distance between sites of that length in a random sequence.
Two validation checks were applied: all 30 recognition sequences must be their own reverse complement (palindrome), 30/30 pass; end types were spot-checked against accepted references for 8 enzymes, all matched.
Practical points to watch out for
The same recognition sequence does not imply the same cut site. SmaI and XmaI both recognize CCCGGG, but SmaI cuts at CCC^GGG giving blunt ends, while XmaI cuts at C^CCGGG leaving a 4 nt 5′ overhang — the products cannot be ligated to each other.
This relationship is called neoschizomers. Isoschizomers, by contrast, share both the recognition sequence and cut site — HpaII and MspI are a classic example.
Average spacing is only an order-of-magnitude guide. 4^L assumes all four bases are equally probable and independent. Neither holds in real genomes: GC content deviates from 50%, and CpG dinucleotides are severely underrepresented in vertebrate genomes, making CG-containing sites far sparser than the formula predicts. Treat this column as a rough sense of scale — “a 6-cutter cuts roughly every few kilobases” — not as a quantitative prediction.
Methylation can block digestion. Sensitivity to methylation often differs among isoschizomers; this is precisely why HpaII and MspI are used in pairs to detect CpG methylation. Whether a given enzyme is blocked by Dam, Dcm, or CpG methylation depends on your supplier’s documentation — this table contains no methylation information.
FAQ
Why do SmaI and XmaI recognize the same sequence but cannot be used interchangeably?
Both recognize CCCGGG, but their cut sites differ: SmaI cuts at CCC^GGG giving blunt ends, while XmaI cuts at C^CCGGG leaving a 4 nt 5′ overhang. Because the product ends are incompatible, they cannot be ligated to each other. This relationship is called neoschizomers. Isoschizomers share both the recognition sequence and cut site — HpaII and MspI are an example.
Can the 'average spacing' column predict how many times an enzyme will cut?
Only as an order-of-magnitude estimate. 4^L assumes all four bases are equally probable and independent, neither of which holds in real genomes — GC content deviates from 50%, and CpG dinucleotides are severely underrepresented in vertebrate genomes, making CG-containing sites far sparser than the formula predicts. To know how many times an enzyme cuts your specific sequence, use the site-finder above.
Why does the site finder not support degenerate bases?
All 30 enzymes in this table have recognition sequences composed of unambiguous bases (ACGT) only — no degenerate positions. If your sequence contains N or other ambiguous codes, expand them to unambiguous form first; otherwise it is impossible to determine whether a given position matches the recognition sequence.
Why do circular and linear sequences give different results?
In circular sequences (plasmids), sites that span the origin are missed in linear mode. Fragment counts also differ: n sites in a linear sequence produce n+1 fragments, whereas n sites in a circular sequence produce n fragments. Choosing the correct topology is important.
Is the data reliable?
Recognition sequences and cut sites are parsed programmatically from the REBASE raw data files (version 609) — no manual transcription. Two validation checks were applied: all 30 recognition sequences must be their own reverse complement (palindrome), 30/30 pass; end types were spot-checked against accepted references for 8 enzymes, all matched.
Related tools
Non-Standard Genetic Code Tables: Mitochondrial, Bacterial, and Ciliate Codon Differences (NCBI transl_table Reference)
All 27 NCBI genetic code tables with per-codon differences from the standard table, searchable by table number or codon.
Codon Usage Table & Rare Codon Scanner (with CAI Calculation)
Codon usage frequencies for nine organisms — paste a CDS to scan for rare codons, compute CAI, and identify the optimal codon per amino acid in your expression host.
DNA / RNA Reverse Complement Online Converter
Paste a sequence to get its reverse complement, complement, or reverse; supports FASTA input and IUPAC degenerate bases.
Buy me a coffee