Translating a sequence is table lookup, and anyone can be handed the table. The marks are in the other three things: which reading frame you are actually in, finding the open reading frame across all six frames rather than just the obvious one, and GC3 as the measure of codon bias. This does all of it from one paste.
Translating a sequence is a lookup table, which is the easy part. What earns marks is the reading frame, the open reading frame across all six of them, and GC3 as the measure of codon bias. This does all of it, and shows the peptide both ways round.
A T or a U both read as thymine, so mRNA and DNA can go in the same box. Spaces, line breaks and lowercase are stripped. All six reading frames are searched, so an ORF on the reverse strand is found too.
Reading frames, and why there are six
The genetic code is read three bases at a time, so a sequence can be partitioned into codons in three different ways depending on where you start. Those are the three forward reading frames. Add the three frames of the reverse complement and you have six, because a gene on the opposite strand is a real possibility and you cannot tell from the sequence alone.
Frame 1 starts at base 1, frame 2 at base 2, frame 3 at base 3. A sequence of n bases gives a complete codon only if (n − offset) is divisible by three, so the last one or two bases of a frame are often incomplete and ignored.
Finding the open reading frame
An open reading frame is the stretch from a start codon to the next in-frame stop codon, and it is the only part of a sequence you can actually call a protein-coding gene. The definition is fussy in a way that matters:
- The start codon is ATG (or AUG in mRNA), which codes for methionine. Any ATG will do to begin scanning, but the one that is the real start is usually the first one after any upstream stop.
- The stop codons are TAA, TAG and TGA. None of them codes for an amino acid, so the ORF ends at the first one in the same frame.
- An ATG with no downstream in-frame stop is a partial ORF. It may be a real gene whose remainder you have not sequenced, which is a very different situation from a complete short peptide.
A common mistake is treating any three bases as a codon without checking the frame, which produces garbage. Another is stopping at the first TAA you see anywhere rather than the first one in your chosen frame, which cuts the ORF short at the wrong place.
For bacteria you also have to allow GTG and TTG as alternative start codons, which is worth knowing but which the tool does not treat as starts by default, since the eukaryotic convention is ATG alone.
Codon bias and GC3
Codon bias is the fact that different codons for the same amino acid are not used equally. It exists because the genetic code is redundant, and the redundancy is not neutral: tRNA abundance, codon–anticodon pairing strength, and RNA structure all push usage towards some codons and away from others.
The cheap way to measure it, and the one every first-year course uses, is GC3: the GC percentage counted only across third codon positions. It is called the wobble position, and it is where the bias concentrates.
- GC3 clearly above the overall GC points to a GC-rich organism. Bacteria sit around 50–60% overall with GC3 near 90% in highly expressed genes.
- GC3 clearly below the overall GC is the usual pattern in most organisms, where third positions are strongly A- and T-biased.
The gap between the two numbers is the signal, which is why this tool reports both and the difference rather than just overall GC.
Why it matters: rare codons slow translation down, so a coding sequence full of them can express poorly in a heterologous host. That is the whole basis of codon optimisation for recombinant expression, and it is also why so-called synonymous substitutions, which change a codon without changing the amino acid, turn out not to be neutral in evolution.
Worked example
Take ATGGCTTGATAA: ATG, GCT, TGA, TAA.
Frame 1 reads ATG = Met, GCT = Ala, TGA = stop. Translation is MA*, and the ORF is 2 codons long and complete. The third and fourth codons, TGA and TAA, are both stop codons and are never translated.
Now swap in AUGGCUUGA and you get exactly the same peptide, because the genetic code does not care whether you wrote U or T. Same answer, which is the point: mRNA and DNA coding sequences can go in the same box.
For codon bias, take GCCGCTGCAGCA. Overall GC is 75%, but the third bases are C, T, A and A, so GC3 is only 25%. That 50-point gap is a strongly A/T-biased coding sequence, and it is the sort of thing that tells you immediately that this is not a bacterial gene.
Things this cannot tell you
A codon table is a dictionary, not a predictor. It knows nothing about whether a reading frame is actually used in a living cell, and that depends on promoters, start codon context, RNA structure, and whether the organism even has the tRNA to read that codon. Equally, a real annotation pipeline uses six-frame translation plus similarity searches against known proteins, not a codon table. If you need to know what a sequence encodes rather than what it could encode, that is a database question and this is not the tool for it.
All of this is the standard genetic code, which also has a small number of exceptions in real organisms. Those are rare, organism-specific, and not something a generic table can represent.
Frequently Asked Questions
How many reading frames are there and why?
Six. Three ways to partition a sequence into codons on the forward strand, and three more on the reverse complement, because a gene on the opposite strand is just as possible and the sequence alone cannot tell you which.
What makes an open reading frame?
It starts at a start codon, ATG, and ends at the first in-frame stop codon, which is TAA, TAG or TGA. Both conditions matter: a stop codon in a different frame does not end the ORF, and an ATG with no downstream in-frame stop is only a partial ORF.
Is a stop codon translated into an amino acid?
No. The three stop codons code for nothing and end translation. They are shown as a star in the peptide so you can see where the ORF stopped rather than silently dropping the tail.
What is GC3 and why is it different from overall GC?
GC3 is the GC percentage counted only across third codon positions, the wobble position where codon bias concentrates. Overall GC covers all three positions, so a big gap between the two numbers is the signal that the sequence is codon biased.
Why does mRNA with U give the same answer as DNA with T?
The genetic code is the same for both. Uracil replaces thymine in RNA and pairs with adenine in exactly the same way, so the two sequences are read identically.
Why is codon bias not a neutral quirk?
Rare codons slow translation down because their tRNAs are less abundant, so a coding sequence full of rare codons can express poorly in a host that does not favour them. That is the basis of codon optimisation for recombinant expression.
Can this tell me what my gene actually encodes?
No. A codon table is a dictionary, not a predictor. Whether a reading frame is used in a living cell depends on promoters, start codon context and RNA structure, and real annotation pipelines search databases of known proteins rather than reading a table.
