Free tools. No account. Nothing stored.

Restriction Enzyme Cut Site Finder

Knowing a DNA sequence and being able to cut it are different skills. Restriction enzymes cut at specific short recognition sequences, and once you know where they cut, the sizes of the pieces are just the distances between those cuts. Get the topology wrong and the answer is wrong, which is the single most common way this question loses marks.

Restriction mapping is the difference between knowing a sequence and being able to cut it. Paste a sequence, pick an enzyme, and this gives you every recognition site plus the actual fragment sizes you would see on a gel. Whether the DNA is linear or circular changes the answer, and that is where most marks are lost.

Spaces, line breaks and lowercase are stripped out automatically. Pick circular for a plasmid: a circle has no ends, so you lose the two terminal fragments a linear molecule keeps, and a site straddling the origin still counts. SmaI cuts blunt in the middle, so it gives one flat end instead of a sticky one.

Linear or circular: the part that catches people

This is worth getting straight before anything else, because a linear molecule and a circle behave differently in two separate ways.

Linear DNA has ends. Those ends are fragment sizes of their own, so a linear molecule cut twice gives three pieces, not two. A 12 bp sequence with EcoRI sites after base 1 and base 7 gives fragments of 1, 6 and 5 bp — and yes, that adds up to 12.

Circular DNA has no ends. Cut the same 12 bp molecule as a circle and you get two pieces of 6 and 6 bp, because there is no start and no end, only the distance all the way round between consecutive cuts. This is why a plasmid cut once is still one linear fragment of its full length, and why a plasmid cut twice gives two fragments rather than three.

A second consequence of circularity: a recognition site that straddles the origin is a real cut even though it spans the end of the sequence you typed. A linear scan will miss it completely. This tool checks every starting position on a circular molecule so it does not.

Sticky ends and blunt ends

Different enzymes cut at different offsets within their recognition site, and that offset is what decides the shape of the ends.

  • EcoRI cuts G^AATTC, leaving a 5′ AATT overhang.
  • SmaI cuts CCC^GGG, straight down the middle, leaving blunt ends.
  • NotI cuts GC^GGCCGC, leaving a 5′ GGCC overhang.
  • KpnI cuts GGTAC^C and SacI cuts GAGCT^C, both leaving 3′ overhangs.

Sticky ends are what make restriction cloning work at all: two fragments cut by the same enzyme have complementary overhangs, so they anneal to each other and the ligase seals the nicks. Blunt ends have nothing to hold onto, so ligation is much less efficient.

Degenerate recognition sites

Some enzymes do not recognise one fixed sequence. They use the IUPAC degenerate codes:

  • R = A or G, Y = C or T
  • S = C or G, W = A or T
  • K = G or T, M = A or C
  • N = any base

So MstII recognises CCTNAGG, where the fourth base can be anything, and ApoI recognises RAATTY. When you count sites for a degenerate enzyme, expect a range rather than a single number, and this tool flags the result so you do not quote one concrete sequence as though it were the whole rule.

Four-cutters like TaqI (TCGA) are worth knowing about for a different reason. They cut roughly every 256 bases on average, which makes them useful for generating random fragments for sequencing but useless for mapping.

How to read the fragment sizes

The sizes are base pairs, and on a gel the largest fragment runs shortest and the smallest runs furthest. Two things routinely trip people up:

  • Equal sizes give one band. A plasmid with three identical 1000 bp fragments shows a single thick band, not three. The number of visible bands is the number of distinct sizes, which is at most the number of fragments.
  • Fragments too small to see. Anything under roughly 100 bp usually runs off the bottom of a standard agarose gel, and anything above about 20 kb needs a different gel or a pulsed-field method. A predicted 30 bp fragment is real and probably invisible.

Worked example

Take GAATTCGAATTC, 12 bp, and digest with EcoRI as a linear fragment.

Both GAATTC sites match. EcoRI cuts one base into the site, so the cuts fall after bases 1 and 7. The fragments are 1 bp, 6 bp and 5 bp.

Now the same sequence as a circular plasmid. There are no ends, so the pieces are the distance from cut 1 to cut 7, which is 6 bp, and the distance from cut 7 all the way round to cut 1, which is also 6 bp. Two fragments, not three.

That is the whole difference, and it is a free mark if you remember which one you are being asked about.

Limitations worth knowing

This tool works on a single strand as you type it. It does not model methylation, and that matters more than it sounds: bacterial methylation at a subset of recognition sites, which is exactly how the restriction-modification system defends a bacterium against its own phage DNA, will block cleavage at those sites. A site present in the sequence is not always a site that gets cut in a real digest. Star activity, where an enzyme cuts slightly off-target under high glycerol or low salt, is the other caveat. Both are lab realities rather than things a sequence calculation can know about.

Frequently Asked Questions

How many fragments do I get from cutting circular DNA twice?

Two, not three. A circle has no ends, so there is no start piece and no end piece, only the distance all the way around between consecutive cuts. Linear DNA cut twice gives three.

Why does it matter whether the DNA is linear or circular?

It changes both the number of fragments and their sizes, because a linear molecule keeps the two pieces at its ends while a circle does not. It is the most common way this question loses marks.

What is the difference between a sticky end and a blunt end?

A sticky end is a short single-stranded overhang left because the enzyme cut off-centre, like EcoRI leaving AATT. A blunt end is a flat cut straight through both strands, like SmaI. Sticky ends anneal to each other, which is how cloning works.

What do the IUPAC letters mean in a recognition site?

They stand for sets of bases. R is A or G, Y is C or T, N is any base, and so on. MstII recognises CCTNAGG with any fourth base, and ApoI recognises RAATTY.

Why does my linear search miss a site near the end of the sequence?

On a circular molecule a site that straddles the origin is a genuine cut, so scanning only positions that fit entirely inside the sequence will miss it. A circular search has to check every starting position.

Are all fragments visible on a gel?

No. Fragments of the same size run as a single thick band, and fragments below roughly 100 bp usually run off the bottom of a standard agarose gel.

Does a recognition site in the sequence always mean the enzyme cuts it?

Not always. Bacterial methylation can block cleavage at a subset of sites, and star activity causes off-target cutting under unusual conditions. A sequence calculation cannot know about either of those.

Related tools