Bench Tools Free browser-based calculators for the lab and for bioinformatics

ToolsSequence tools

Codon Usage Table & Rare Codon Scanner (with CAI Calculation)

Codon usage frequencies for nine organisms — paste a CDS to scan for rare codons, compute CAI, and identify the optimal codon per amino acid in your expression host.

Codon usage frequencies for nine commonly used organisms, plus a rare codon scanner: paste a CDS sequence, select your expression host, and immediately see which positions will slow down translation.

Organisms and sample sizes

Organism CDS count Total codons Reliable?
Escherichia coli W3110 4,332 1,372,057
Homo sapiens 93,487 40,662,582
Mus musculus 53,036 24,533,776
Saccharomyces cerevisiae 14,411 6,534,504
Cricetulus griseus 331 153,527 ⚠ small sample
Drosophila melanogaster 42,417 21,945,319
Caenorhabditis elegans 24,994 11,197,796
Danio rerio 19,062 8,042,248
Arabidopsis thaliana 80,395 31,098,475

Codon usage: E. coli W3110

AA Codon Fraction per 1000 Family optimum
Stop TAA 0.64 2.0
Stop TAG 0.07 0.2
Stop TGA 0.29 0.9
A GCA 0.21 20.1
A GCC 0.27 25.7
A GCG 0.36 33.9
A GCT 0.16 15.2
C TGC 0.56 6.4
C TGT 0.44 5.1
D GAC 0.37 19.1
D GAT 0.63 32.2
E GAA 0.69 39.7
E GAG 0.31 18.0
F TTC 0.43 16.5
F TTT 0.57 22.2
G GGA 0.11 7.9
G GGC 0.41 29.8
G GGG 0.15 11.0
G GGT 0.34 24.7
H CAC 0.43 9.8
H CAT 0.57 13.0
I ATA 0.07 4.2
I ATC 0.42 25.2
I ATT 0.51 30.4
K AAA 0.76 33.6
K AAG 0.24 10.3
L CTA 0.04 3.8
L CTC 0.10 11.1
L CTG 0.50 53.1
L CTT 0.10 11.0
L TTA 0.13 13.8
L TTG 0.13 13.6
M ATG 1.00 27.8
N AAC 0.55 21.6
N AAT 0.45 17.6
P CCA 0.19 8.4
P CCC 0.12 5.5
P CCG 0.53 23.4
P CCT 0.16 7.0
Q CAA 0.35 15.4
Q CAG 0.65 29.0
R AGA 0.04 2.0
R AGG 0.02 1.1
R CGA 0.06 3.5
R CGC 0.40 22.3
R CGG 0.10 5.4
R CGT 0.38 21.0
S AGC 0.28 16.1
S AGT 0.15 8.7
S TCA 0.12 7.0
S TCC 0.15 8.6
S TCG 0.15 8.9
S TCT 0.15 8.4
T ACA 0.13 6.9
T ACC 0.44 23.5
T ACG 0.27 14.4
T ACT 0.16 8.8
V GTA 0.15 10.9
V GTC 0.22 15.3
V GTG 0.37 26.3
V GTT 0.26 18.2
W TGG 1.00 15.2
Y TAC 0.43 12.2
Y TAT 0.57 16.1

Codon usage: Homo sapiens

AA Codon Fraction per 1000 Family optimum
Stop TAA 0.30 1.0
Stop TAG 0.24 0.8
Stop TGA 0.47 1.6
A GCA 0.23 15.8
A GCC 0.40 27.7
A GCG 0.11 7.4
A GCT 0.27 18.4
C TGC 0.54 12.6
C TGT 0.46 10.6
D GAC 0.54 25.1
D GAT 0.46 21.8
E GAA 0.42 29.0
E GAG 0.58 39.6
F TTC 0.54 20.3
F TTT 0.46 17.6
G GGA 0.25 16.5
G GGC 0.34 22.2
G GGG 0.25 16.5
G GGT 0.16 10.8
H CAC 0.58 15.1
H CAT 0.42 10.9
I ATA 0.17 7.5
I ATC 0.47 20.8
I ATT 0.36 16.0
K AAA 0.43 24.4
K AAG 0.57 31.9
L CTA 0.07 7.2
L CTC 0.20 19.6
L CTG 0.40 39.6
L CTT 0.13 13.2
L TTA 0.08 7.7
L TTG 0.13 12.9
M ATG 1.00 22.0
N AAC 0.53 19.1
N AAT 0.47 17.0
P CCA 0.28 16.9
P CCC 0.32 19.8
P CCG 0.11 6.9
P CCT 0.29 17.5
Q CAA 0.27 12.3
Q CAG 0.73 34.2
R AGA 0.21 12.2
R AGG 0.21 12.0
R CGA 0.11 6.2
R CGC 0.18 10.4
R CGG 0.20 11.4
R CGT 0.08 4.5
S AGC 0.24 19.5
S AGT 0.15 12.1
S TCA 0.15 12.2
S TCC 0.22 17.7
S TCG 0.05 4.4
S TCT 0.19 15.2
T ACA 0.28 15.1
T ACC 0.36 18.9
T ACG 0.11 6.1
T ACT 0.25 13.1
V GTA 0.12 7.1
V GTC 0.24 14.5
V GTG 0.46 28.1
V GTT 0.18 11.0
W TGG 1.00 13.2
Y TAC 0.56 15.3
Y TAT 0.44 12.2

Check the sample size before trusting the numbers

This is the most important point on this page. Codon usage tables are derived by counting codons across a set of genes — the fewer genes in that set, the less trustworthy the table. Most codon usage pages give you numbers without telling you the sample size.

The Kazusa database entry for “Escherichia coli K12” is based on only 14 CDS. E. coli is by far the most commonly used host for codon optimization, yet using a table derived from 14 genes to optimize a sequence is like inferring a country’s height distribution from three people. This page uses W3110 (4,332 CDS), the standard K-12 laboratory strain with genome-level coverage.

Every organism in the table above shows its CDS count; entries with fewer than 1,000 CDS are flagged as “small sample” — CHO cells have only 331 CDS, which matters even though CHO is a widely used protein expression host.

Why rare codons matter

Synonymous codons differ dramatically in their corresponding tRNA abundance. When a sequence contains a run of codons that are rare in the host, the ribosome stalls there, potentially reducing yield, causing premature termination, or inducing frameshifting. One of the most common causes of failed heterologous expression is inserting a eukaryotic gene rich in AGA/AGG (rare arginine codons in E. coli) directly into E. coli.

The scanner flags codons whose relative frequency is below 0.10 and shows the optimal codon for that amino acid in the selected host.

About CAI

The Codon Adaptation Index (CAI) was introduced by Sharp and Li (Nucleic Acids Res. 1987;15(3):1281–1295). It is defined as the geometric mean of the relative adaptiveness w of each codon, where w is the usage frequency of that codon divided by the frequency of the most common codon for the same amino acid family. A CAI closer to 1 indicates that the sequence codon usage more closely matches host preference.

One important distinction must be made clear: Sharp and Li’s original formulation uses a high-expression gene reference set to compute w, whereas this page uses the genome-wide average. This is a common simplification, but it introduces a systematic offset relative to CAI values computed from high-expression reference sets, so the numbers here cannot be directly compared with CAI values reported in the literature. They are appropriate for comparing different sequence variants against the same host.

Methionine, tryptophan (each encoded by a single codon, so w is always 1), and stop codons are excluded from the calculation.

FAQ

Why not use the 'Escherichia coli K12' entry from Kazusa?

Because it is based on only 14 CDS. Codon usage bias is a statistical measure, and 14 genes are not enough to produce a reliable distribution. This page uses W3110 (4,332 CDS), the standard K-12 laboratory strain with genome-level coverage. The first thing to check when reading a codon table is how many sequences it is based on — most pages don't tell you.

Can the CAI calculated here be compared directly with values in published papers?

No. Sharp and Li's original definition uses a high-expression gene reference set to compute relative adaptiveness w; this page uses the genome-wide average instead. This is a common simplification, but it introduces a systematic offset. The values here are appropriate for comparing different sequence variants against the same host, not for cross-referencing absolute CAI values from the literature.

What threshold should I use for rare codons?

The default of 0.10 is a common starting point. What matters most is not individual rare codons but **runs of consecutive rare codons** — isolated occurrences are usually harmless, whereas clusters are what tend to cause ribosome pausing. The scan results will specifically flag adjacent rare codons.

Will replacing all rare codons improve expression?

Not necessarily. Codon optimization addresses only the translational rate layer. Failed expression can also stem from secondary structure near the 5′ end of the mRNA, depletion of rare tRNA pools, protein misfolding, or toxicity. Moreover, aggressive optimization (replacing everything with the most frequent codon) can sometimes reduce solubility, because translation that is too fast can disrupt co-translational folding.

How reliable is the CHO cell data?

CHO (Chinese hamster ovary) has only 331 CDS in Kazusa; this page flags it as 'small sample'. It is included because CHO is a widely used protein expression host, but the confidence in those numbers is substantially lower than for human or mouse. Factor that in when interpreting results.

Related tools

DNA / RNA Reverse Complement Online Converter

Paste a sequence to get its reverse complement, complement, or reverse; supports FASTA input and IUPAC degenerate bases.

Online DNA GC Content & Tm Calculator — GC%, Base Composition, Molecular Weight

Calculate GC%, base composition, molecular weight, and two empirical Tm values for rapid primer screening.

DNA Sequence to Protein Translation (Six-Frame) — Online Tool

Standard genetic code (NCBI table 1): single-frame or six-frame translation with ORF detection.

No ads, no tracking, no sign-up — and every formula here is checked against a known answer. Keeping it that way takes ongoing work. If it saved you time, buy me a coffee.
Buy me a coffee