Three questions come up every single time you touch a DNA sequence, and they always take longer by hand than they should. What is the other strand? What does the mRNA say? And what percentage of this is G and C?
This does all three at once, plus molecular weight and a melting temperature estimate, because once you have the GC percentage those are one substitution away. Paste anything in — lowercase, spaces, line breaks all get cleaned up first.
Paste a strand and get the three things you are always asked for: the reverse complement, the mRNA transcript, and the GC content. Molecular weight and melting temperature come along too, because once you have the GC percentage they are one substitution away.
Reading the two directions properly
DNA strands run in opposite directions. We write both of them 5′ to 3′, which is why you cannot just complement a sequence and call it done — you also have to read it backwards.
Take ATGCGCTA:
- Your strand, 5′ → 3′: ATGCGCTA
- Its complement, written antiparallel: TAGCGCAT
- Read that complement 5′ → 3′ and you get the same string back, which is the reverse complement
That double reversal is why the answer is not simply the complement. If you have ever been caught out by this, the trick is to remember that the reverse complement is its own inverse — run the tool on its own output and you get your original sequence back. That is a free check on your working.
A palindrome is a sequence that equals its own reverse complement. ATGCAT and GGATCC both are, which is not a coincidence: palindromic sites are exactly where restriction enzymes bind, cutting both strands at the same point.
Transcription: same letters, T becomes U
Transcription copies the template strand into RNA. The only difference in the alphabet is that RNA uses uracil where DNA uses thymine, because uracil pairs with adenine without the methyl group that makes thymine harder to tell apart from cytosine.
So the mRNA of ATGCGCTA is UACGCGAU. Every other letter carries over unchanged.
Why GC content is the number everyone quotes
GC content is simply (G + C) divided by the total, times 100. For ATGCGCTA that is 4 G or C out of 8, so 50%.
It matters because G–C pairs form three hydrogen bonds while A–T pairs form only two. More GC means a duplex that is harder to pull apart, which means:
- Higher melting temperature, so PCR needs a hotter annealing step
- More resistance to chemicals that break single strands, like alkali
- Higher buoyant density, which is how DNA gets separated by centrifuge
That is also why every commercial primer pair is quoted with a GC percentage, and why a primer at 10% GC behaves nothing like one at 60%.
Molecular weight and melting temperature
Molecular weight uses the standard genomic DNA approximation, where n is the number of base pairs:
MW = n × (249 + 0.98 × %GC)
So a 1,000 bp sequence at 50% GC comes out at 298,000 g/mol. It is an approximation — real bases carry different masses and a strand is never a uniform polymer — but it is the number every textbook quotes.
Melting temperature follows the Marmur relation:
Tm = 81.5 + 0.41(%GC) − 675 ÷ n
One thing this tool will not do: report a melting temperature for anything under 14 base pairs. The 675 divided by n term runs away on very short oligos and returns a large negative number, and a negative melting point is not a physical quantity. Above that length the estimate means something.
Three things that catch people out
A strand is a sequence, not a molecule you can weigh. Molecular weight needs a duplex, so the input length is halved. An 8-base sequence describes a 4 base pair duplex.
Direction is not decoration. The 5′ and 3′ ends are chemically different, which is why replication can only proceed one way along a strand. Losing them is how people end up with the right letters in the wrong order.
Uracil means RNA, and the reverse complement still works. U pairs with A exactly as T does, so the tool handles RNA input without a separate mode. GC content is unaffected either way, since neither U nor T is a G or a C.
Frequently asked questions
How do I find the reverse complement of a DNA sequence?
Complement each base (A becomes T, T becomes A, G becomes C, C becomes G) and then read the result backwards. Equivalently, read the original backwards and then complement it.
What is the reverse complement of ATGCAT?
ATGCAT. Complementing gives TACGTA and reversing that gives ATGCAT back, so it is its own reverse complement — a palindrome.
How is GC content calculated?
Add the G and C counts, divide by the total length, multiply by 100. A sequence with 4 G or C bases out of 8 is 50%.
What does a high GC content mean for PCR?
Higher melting temperature, so the annealing step has to be hotter. It also raises the risk of primers folding back on themselves, because GC-rich sequences form stable internal structures more easily.
Why does mRNA have U instead of T?
Uracil pairs with adenine the same way thymine does, but without the methyl group that lets thymine be mistaken for cytosine. RNA does not get the proofreading DNA does, so the pairing does not need to be quite so fussy.
Is the molecular weight exact?
No. It is a standard approximation that assumes a uniform base composition. For sequencing QC and rough yield estimates it is close enough, and it is the figure lab manuals quote.
Why does a sequence that equals its own reverse complement matter?
It is a palindrome, and palindromic sites are where restriction enzymes bind. Because the site reads the same on both strands, the enzyme recognises it from either direction and cuts both strands at the same base pair.
Related tools
- Punnett Square Calculator — work out the offspring ratios from a cross
- Population Genetics Calculator — Hardy-Weinberg ratios and a chi-square test on real counts
- Population Growth Calculator — exponential and logistic population projection
- All student GPA tools
