Accepts single-letter IUPAC codes. FASTA headers and whitespace are removed automatically.
Introduction
The Amino Acid Counter lets researchers and students instantly analyse the residue composition of any protein sequence. Paste a sequence in single-letter code or upload a FASTA file to get a full breakdown of counts, percentages, and physicochemical group distributions — essential data for purification planning, structural prediction, and comparative proteomics.
About the Tool
The Amino Acid Counter is a free, browser-based calculator that analyses the residue composition of any protein sequence. Enter a sequence in single-letter IUPAC code, or upload a .txt or .fasta file, and the tool reports the exact count and percentage of each of the 20 standard amino acids, plus a breakdown across five physicochemical categories — hydrophobic, polar uncharged, positively charged, negatively charged, and glycine. All calculations run entirely client-side in your browser, so no sequence data is ever transmitted to a server, and no account or login is required.
Input Explanation
Paste your protein sequence in standard single-letter amino acid code into the text area. The tool accepts raw sequences as well as FASTA-formatted input — header lines beginning with > are stripped automatically before counting, and whitespace, numbers, and line breaks are removed so you do not need to clean your sequence first. Alternatively, click Upload .txt / .fasta file to load a sequence directly from disk. A live character counter beneath the field helps you confirm sequence length before running the calculation. Non-standard characters such as B, Z, or X are ignored and reported in the result note rather than silently miscounted.
Formula Explanation
Each amino acid's abundance is calculated with the standard residue frequency equation:
Variables: count of X is the number of times a given amino acid appears in the sequence; total valid residues is the sum of all 20 standard amino acid counts after non-standard or ambiguous characters (B, Z, X) have been excluded.
Units: the result is expressed as a percentage (%) of total valid residues.
Scientific explanation: because the denominator excludes ambiguous characters from both numerator and denominator, percentages across all detected amino acids always sum to 100%. This matches the residue-frequency method used by ExPASy ProtParam. The relative abundance bar shown in the results table is separately normalised to the single most frequent residue, which makes it easy to visually compare rare and dominant residues — but note this bar is not itself a percentage of the full sequence.
Worked Example
Sample Input
The 140-residue example sequence loaded by the "Load Example Sequence" button:
MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQFEVVHSLAKWKRQTLGQHDFSAGEGLYTHMKALRPDEDRLSPLHSVYVDQWDWELVMGDGERAFSTLHYRKGNS
Step-by-Step Calculation
1. The tool strips any FASTA header and whitespace, leaving 140 valid residues (no ambiguous characters present).
2. Each residue is tallied. Leucine (L) occurs most frequently, with 14 occurrences.
3. Leucine's percentage is calculated as: % L = (14 ÷ 140) × 100 = 10.0%. The same formula is applied to all 19 remaining residues.
Final Result
Total residues: 140. Most abundant residue: Leucine at 10.0% (14 occurrences). Physicochemical breakdown: Hydrophobic 39.3%, Polar uncharged 22.9%, Positively charged 17.9%, Negatively charged 12.9%, Glycine 7.1%.
Interpretation
The hydrophobic fraction (39.3%) is notably higher than the combined charged fraction (30.7%), suggesting this sequence would likely retain well on a hydrophobic interaction chromatography (HIC) column. The moderate positive-to-negative charge ratio (17.9% vs 12.9%) indicates a net positive charge at neutral pH, which is also useful for anticipating behaviour on an ion-exchange column.
Result Interpretation
The summary grid at the top of your results shows total residues counted, unique amino acids found, the single-letter code of the most abundant residue with its raw count, and that residue's percentage abundance. The physicochemical group section displays each category's percentage of the full sequence, followed by chips for each detected residue showing count and percentage — categories entirely absent from your sequence are hidden. The full composition table sorts every residue from most to least abundant and includes a relative abundance bar normalised to the top residue. Use the percentage column alongside known functional domains to flag unusual compositions — for example, cysteine content above 5% in a cytoplasmic protein warrants investigation of disulfide bond potential or redox-sensitive active sites.
Practical Applications
Composition analysis is valuable at multiple stages of protein research. Before expression system selection, unusual composition can flag the need for a codon-optimised gene or specialised host strain. Before chromatographic purification, the charged-to-hydrophobic residue ratio helps choose between ion-exchange chromatography (IEX), hydrophobic interaction chromatography (HIC), or affinity methods. When assessing stability, a high hydrophobic fraction in a soluble protein may indicate aggregation-prone regions requiring buffer optimisation. Composition data is also a prerequisite input for calculating theoretical isoelectric point (pI) and extinction coefficient, both available via related tools on this site.
Scientific Notes & Limitations
The tool assigns each of the 20 canonical amino acids to one of five physicochemical categories. Hydrophobic residues (A, V, I, L, M, F, W, P) have non-polar side chains that localise to the protein core and are enriched in membrane proteins and aggregation-prone sequences. Polar uncharged residues (S, T, C, Y, N, Q) can form hydrogen bonds and include common phosphorylation sites (Ser, Thr) and disulfide-forming cysteine. Positively charged residues at neutral pH (K, R, H) are enriched in DNA-binding domains and nuclear localisation signals. Negatively charged residues (D, E) contribute to surface charge, calcium coordination, and catalytic active sites. Glycine occupies a special category due to its exceptional backbone flexibility.
Limitations: ambiguous codes (B, Z, X) are excluded rather than resolved, so results reflect only unambiguous residues. Multi-sequence FASTA input is concatenated into one combined composition rather than analysed per-sequence — split files first if per-sequence results are required. Composition alone does not account for post-translational modifications, glycosylation, or three-dimensional folding, all of which also influence real-world purification and stability behaviour.
Amino Acid Molecular Weight & pKa Reference| Code | Amino Acid | MW (Da) | Side Chain pKa | Category |
|---|---|---|---|---|
| D | Aspartic acid | 115.09 | 3.65 | Negatively charged |
| E | Glutamic acid | 129.12 | 4.25 | Negatively charged |
| H | Histidine | 137.14 | 6.00 | Positively charged |
| C | Cysteine | 121.16 | 8.30 | Polar uncharged |
| Y | Tyrosine | 163.18 | 10.07 | Polar uncharged |
| K | Lysine | 128.17 | 10.53 | Positively charged |
| R | Arginine | 156.19 | 12.48 | Positively charged |
| G | Glycine | 75.07 | — | Special / Glycine |
| A | Alanine | 89.09 | — | Hydrophobic |
| W | Tryptophan | 204.23 | — | Hydrophobic |
Practical Tips
- Compare the combined charged-residue percentage against the hydrophobic percentage before choosing between ion-exchange and hydrophobic interaction chromatography.
- Check the histidine count before deciding whether IMAC purification is viable without adding a His-tag.
- Run the calculator before ordering a synthetic gene — unusual amino acid usage can indicate the need for host-specific codon optimisation.
- Flag any cysteine content above 5% for follow-up investigation of disulfide bond potential in cytoplasmic proteins.
- Use the character counter to sanity-check sequence length against the expected size of your protein before trusting the results.
Common Mistakes to Avoid
A frequent error is pasting nucleotide sequences instead of protein sequences — ATGC characters are not valid amino acid codes and will be flagged in the result note so you can detect this immediately. Another common mistake is expecting ambiguous residue codes such as B (Asp or Asn), Z (Glu or Gln), or X (unknown) to be counted; these are excluded by design to ensure percentages reflect only unambiguous residues. When working with multi-sequence FASTA files, remember all sequences are concatenated after header removal — split the file first if per-sequence composition is required. Finally, do not interpret the relative abundance bar as an absolute scale; it is normalised to the most frequent residue, not to 100% of the sequence.