Codon Usage Table & Rare Codon Scanner (with CAI Calculation)
Codon usage frequencies for nine commonly used organisms, plus a rare codon scanner: paste a CDS sequence, select your expression host, and immediately see which positions will slow down translation.
Organisms and sample sizes
| Organism | CDS count | Total codons | Reliable? |
|---|---|---|---|
| Escherichia coli W3110 | 4,332 | 1,372,057 | ✓ |
| Homo sapiens | 93,487 | 40,662,582 | ✓ |
| Mus musculus | 53,036 | 24,533,776 | ✓ |
| Saccharomyces cerevisiae | 14,411 | 6,534,504 | ✓ |
| Cricetulus griseus | 331 | 153,527 | ⚠ small sample |
| Drosophila melanogaster | 42,417 | 21,945,319 | ✓ |
| Caenorhabditis elegans | 24,994 | 11,197,796 | ✓ |
| Danio rerio | 19,062 | 8,042,248 | ✓ |
| Arabidopsis thaliana | 80,395 | 31,098,475 | ✓ |
Codon usage: E. coli W3110
| AA | Codon | Fraction | per 1000 | Family optimum |
|---|---|---|---|---|
| Stop | TAA |
0.64 | 2.0 | ★ |
| Stop | TAG |
0.07 | 0.2 | |
| Stop | TGA |
0.29 | 0.9 | |
| A | GCA |
0.21 | 20.1 | |
| A | GCC |
0.27 | 25.7 | |
| A | GCG |
0.36 | 33.9 | ★ |
| A | GCT |
0.16 | 15.2 | |
| C | TGC |
0.56 | 6.4 | ★ |
| C | TGT |
0.44 | 5.1 | |
| D | GAC |
0.37 | 19.1 | |
| D | GAT |
0.63 | 32.2 | ★ |
| E | GAA |
0.69 | 39.7 | ★ |
| E | GAG |
0.31 | 18.0 | |
| F | TTC |
0.43 | 16.5 | |
| F | TTT |
0.57 | 22.2 | ★ |
| G | GGA |
0.11 | 7.9 | |
| G | GGC |
0.41 | 29.8 | ★ |
| G | GGG |
0.15 | 11.0 | |
| G | GGT |
0.34 | 24.7 | |
| H | CAC |
0.43 | 9.8 | |
| H | CAT |
0.57 | 13.0 | ★ |
| I | ATA |
0.07 | 4.2 | |
| I | ATC |
0.42 | 25.2 | |
| I | ATT |
0.51 | 30.4 | ★ |
| K | AAA |
0.76 | 33.6 | ★ |
| K | AAG |
0.24 | 10.3 | |
| L | CTA |
0.04 | 3.8 | |
| L | CTC |
0.10 | 11.1 | |
| L | CTG |
0.50 | 53.1 | ★ |
| L | CTT |
0.10 | 11.0 | |
| L | TTA |
0.13 | 13.8 | |
| L | TTG |
0.13 | 13.6 | |
| M | ATG |
1.00 | 27.8 | ★ |
| N | AAC |
0.55 | 21.6 | ★ |
| N | AAT |
0.45 | 17.6 | |
| P | CCA |
0.19 | 8.4 | |
| P | CCC |
0.12 | 5.5 | |
| P | CCG |
0.53 | 23.4 | ★ |
| P | CCT |
0.16 | 7.0 | |
| Q | CAA |
0.35 | 15.4 | |
| Q | CAG |
0.65 | 29.0 | ★ |
| R | AGA |
0.04 | 2.0 | |
| R | AGG |
0.02 | 1.1 | |
| R | CGA |
0.06 | 3.5 | |
| R | CGC |
0.40 | 22.3 | ★ |
| R | CGG |
0.10 | 5.4 | |
| R | CGT |
0.38 | 21.0 | |
| S | AGC |
0.28 | 16.1 | ★ |
| S | AGT |
0.15 | 8.7 | |
| S | TCA |
0.12 | 7.0 | |
| S | TCC |
0.15 | 8.6 | |
| S | TCG |
0.15 | 8.9 | |
| S | TCT |
0.15 | 8.4 | |
| T | ACA |
0.13 | 6.9 | |
| T | ACC |
0.44 | 23.5 | ★ |
| T | ACG |
0.27 | 14.4 | |
| T | ACT |
0.16 | 8.8 | |
| V | GTA |
0.15 | 10.9 | |
| V | GTC |
0.22 | 15.3 | |
| V | GTG |
0.37 | 26.3 | ★ |
| V | GTT |
0.26 | 18.2 | |
| W | TGG |
1.00 | 15.2 | ★ |
| Y | TAC |
0.43 | 12.2 | |
| Y | TAT |
0.57 | 16.1 | ★ |
Codon usage: Homo sapiens
| AA | Codon | Fraction | per 1000 | Family optimum |
|---|---|---|---|---|
| Stop | TAA |
0.30 | 1.0 | |
| Stop | TAG |
0.24 | 0.8 | |
| Stop | TGA |
0.47 | 1.6 | ★ |
| A | GCA |
0.23 | 15.8 | |
| A | GCC |
0.40 | 27.7 | ★ |
| A | GCG |
0.11 | 7.4 | |
| A | GCT |
0.27 | 18.4 | |
| C | TGC |
0.54 | 12.6 | ★ |
| C | TGT |
0.46 | 10.6 | |
| D | GAC |
0.54 | 25.1 | ★ |
| D | GAT |
0.46 | 21.8 | |
| E | GAA |
0.42 | 29.0 | |
| E | GAG |
0.58 | 39.6 | ★ |
| F | TTC |
0.54 | 20.3 | ★ |
| F | TTT |
0.46 | 17.6 | |
| G | GGA |
0.25 | 16.5 | |
| G | GGC |
0.34 | 22.2 | ★ |
| G | GGG |
0.25 | 16.5 | |
| G | GGT |
0.16 | 10.8 | |
| H | CAC |
0.58 | 15.1 | ★ |
| H | CAT |
0.42 | 10.9 | |
| I | ATA |
0.17 | 7.5 | |
| I | ATC |
0.47 | 20.8 | ★ |
| I | ATT |
0.36 | 16.0 | |
| K | AAA |
0.43 | 24.4 | |
| K | AAG |
0.57 | 31.9 | ★ |
| L | CTA |
0.07 | 7.2 | |
| L | CTC |
0.20 | 19.6 | |
| L | CTG |
0.40 | 39.6 | ★ |
| L | CTT |
0.13 | 13.2 | |
| L | TTA |
0.08 | 7.7 | |
| L | TTG |
0.13 | 12.9 | |
| M | ATG |
1.00 | 22.0 | ★ |
| N | AAC |
0.53 | 19.1 | ★ |
| N | AAT |
0.47 | 17.0 | |
| P | CCA |
0.28 | 16.9 | |
| P | CCC |
0.32 | 19.8 | ★ |
| P | CCG |
0.11 | 6.9 | |
| P | CCT |
0.29 | 17.5 | |
| Q | CAA |
0.27 | 12.3 | |
| Q | CAG |
0.73 | 34.2 | ★ |
| R | AGA |
0.21 | 12.2 | ★ |
| R | AGG |
0.21 | 12.0 | |
| R | CGA |
0.11 | 6.2 | |
| R | CGC |
0.18 | 10.4 | |
| R | CGG |
0.20 | 11.4 | |
| R | CGT |
0.08 | 4.5 | |
| S | AGC |
0.24 | 19.5 | ★ |
| S | AGT |
0.15 | 12.1 | |
| S | TCA |
0.15 | 12.2 | |
| S | TCC |
0.22 | 17.7 | |
| S | TCG |
0.05 | 4.4 | |
| S | TCT |
0.19 | 15.2 | |
| T | ACA |
0.28 | 15.1 | |
| T | ACC |
0.36 | 18.9 | ★ |
| T | ACG |
0.11 | 6.1 | |
| T | ACT |
0.25 | 13.1 | |
| V | GTA |
0.12 | 7.1 | |
| V | GTC |
0.24 | 14.5 | |
| V | GTG |
0.46 | 28.1 | ★ |
| V | GTT |
0.18 | 11.0 | |
| W | TGG |
1.00 | 13.2 | ★ |
| Y | TAC |
0.56 | 15.3 | ★ |
| Y | TAT |
0.44 | 12.2 |
Check the sample size before trusting the numbers
This is the most important point on this page. Codon usage tables are derived by counting codons across a set of genes — the fewer genes in that set, the less trustworthy the table. Most codon usage pages give you numbers without telling you the sample size.
The Kazusa database entry for “Escherichia coli K12” is based on only 14 CDS. E. coli is by far the most commonly used host for codon optimization, yet using a table derived from 14 genes to optimize a sequence is like inferring a country’s height distribution from three people. This page uses W3110 (4,332 CDS), the standard K-12 laboratory strain with genome-level coverage.
Every organism in the table above shows its CDS count; entries with fewer than 1,000 CDS are flagged as “small sample” — CHO cells have only 331 CDS, which matters even though CHO is a widely used protein expression host.
Why rare codons matter
Synonymous codons differ dramatically in their corresponding tRNA abundance. When a sequence contains a run of codons that are rare in the host, the ribosome stalls there, potentially reducing yield, causing premature termination, or inducing frameshifting. One of the most common causes of failed heterologous expression is inserting a eukaryotic gene rich in AGA/AGG (rare arginine codons in E. coli) directly into E. coli.
The scanner flags codons whose relative frequency is below 0.10 and shows the optimal codon for that amino acid in the selected host.
About CAI
The Codon Adaptation Index (CAI) was introduced by Sharp and Li (Nucleic Acids Res. 1987;15(3):1281–1295). It is defined as the geometric mean of the relative adaptiveness w of each codon, where w is the usage frequency of that codon divided by the frequency of the most common codon for the same amino acid family. A CAI closer to 1 indicates that the sequence codon usage more closely matches host preference.
One important distinction must be made clear: Sharp and Li’s original formulation uses a high-expression gene reference set to compute w, whereas this page uses the genome-wide average. This is a common simplification, but it introduces a systematic offset relative to CAI values computed from high-expression reference sets, so the numbers here cannot be directly compared with CAI values reported in the literature. They are appropriate for comparing different sequence variants against the same host.
Methionine, tryptophan (each encoded by a single codon, so w is always 1), and stop codons are excluded from the calculation.
FAQ
Why not use the 'Escherichia coli K12' entry from Kazusa?
Because it is based on only 14 CDS. Codon usage bias is a statistical measure, and 14 genes are not enough to produce a reliable distribution. This page uses W3110 (4,332 CDS), the standard K-12 laboratory strain with genome-level coverage. The first thing to check when reading a codon table is how many sequences it is based on — most pages don't tell you.
Can the CAI calculated here be compared directly with values in published papers?
No. Sharp and Li's original definition uses a high-expression gene reference set to compute relative adaptiveness w; this page uses the genome-wide average instead. This is a common simplification, but it introduces a systematic offset. The values here are appropriate for comparing different sequence variants against the same host, not for cross-referencing absolute CAI values from the literature.
What threshold should I use for rare codons?
The default of 0.10 is a common starting point. What matters most is not individual rare codons but **runs of consecutive rare codons** — isolated occurrences are usually harmless, whereas clusters are what tend to cause ribosome pausing. The scan results will specifically flag adjacent rare codons.
Will replacing all rare codons improve expression?
Not necessarily. Codon optimization addresses only the translational rate layer. Failed expression can also stem from secondary structure near the 5′ end of the mRNA, depletion of rare tRNA pools, protein misfolding, or toxicity. Moreover, aggressive optimization (replacing everything with the most frequent codon) can sometimes reduce solubility, because translation that is too fast can disrupt co-translational folding.
How reliable is the CHO cell data?
CHO (Chinese hamster ovary) has only 331 CDS in Kazusa; this page flags it as 'small sample'. It is included because CHO is a widely used protein expression host, but the confidence in those numbers is substantially lower than for human or mouse. Factor that in when interpreting results.
Related tools
DNA / RNA Reverse Complement Online Converter
Paste a sequence to get its reverse complement, complement, or reverse; supports FASTA input and IUPAC degenerate bases.
Online DNA GC Content & Tm Calculator — GC%, Base Composition, Molecular Weight
Calculate GC%, base composition, molecular weight, and two empirical Tm values for rapid primer screening.
DNA Sequence to Protein Translation (Six-Frame) — Online Tool
Standard genetic code (NCBI table 1): single-frame or six-frame translation with ORF detection.
Buy me a coffee