A chemical substance rarely has one identifier. It has a CAS number, an EC number, an InChIKey, a SMILES string and a handful of database keys, and each was created to solve a different problem. Knowing which one to quote — and which one a reader can verify — is a practical skill, not a formality.
EC number
An EC number identifies a substance in the European Community inventory. It takes the form 200-753-7: three digits, three digits, and a check digit.
Three legacy lists feed it, and the range tells you which:
- EINECS — substances on the Community market between 1971 and 1981. Numbers begin 2 or 3.
- ELINCS — substances notified as new after 1981. Numbers begin 4.
- NLP — no-longer-polymers, a reclassification exercise. Numbers begin 5.
The EC number is the identifier European regulation is written around. Annex VI entries carry it, inventory lists key on it, and a document intended for an EU reader should quote it where one exists. Not every substance has one.
InChI and InChIKey
InChI — the IUPAC International Chemical Identifier — is derived from the structure by a published algorithm. Two people with the same structure produce the same InChI without consulting any registry, which is the property CAS and EC numbers lack.
An InChI string is long and layered. The InChIKey is a fixed-length hash of it, in the form UHOVQNZJYSORNB-UHFFFAOYSA-N: fourteen characters for the skeleton, ten for stereochemistry and isotopes, and a final character for the protonation state.
The practical consequence is useful: two InChIKeys that share the first fourteen characters describe the same molecular skeleton, differing only in stereochemistry or charge. That comparison can be made by eye.
SMILES
SMILES is a line notation for a structure — benzene is c1ccccc1. It is compact and human-readable, and it is the format most software accepts for drawing or searching a structure.
It is not canonical by default. The same molecule can be written as several valid SMILES strings, so two differing strings do not prove two different substances. Canonical SMILES algorithms exist, but a string quoted in a document rarely states which was used.
Database keys
These identify a record in one database rather than a substance in the abstract:
- PubChem CID — the NCBI compound record.
- ChemSpider ID — the Royal Society of Chemistry database.
- ChEMBL ID — bioactivity data.
- UNII — the FDA’s unique ingredient identifier.
- KEGG, DrugBank, MeSH, HMDB, NSC — metabolic pathways, drug data, medical indexing, human metabolites, and the NCI’s historical compound numbering.
- Wikidata Q-number — a cross-reference hub, useful for moving between the others.
Which to quote, and when
| Purpose | Identifier |
|---|---|
| EU regulatory document | CAS and EC number |
| Structure search or software input | InChI, InChIKey or SMILES |
| Reaching a body of data | the relevant database key |
| Proving two records are the same substance | InChIKey, then CAS |
Why a document should show several
An identifier can be wrong. A CAS number can be transcribed incorrectly, and a plausible-looking one can be invented — though a wrong CAS number can usually be caught by its check digit. A structure-derived identifier alongside a registry identifier lets a reader confirm that both describe the same thing.
Coverage also tells you something on its own. A substance resolved across thirteen identifier systems is well characterised and easy to check. One resolved across two is not, and a document should say which situation applies rather than presenting both the same way.
Every MolGod substance page resolves identity across sixteen systems and states how many were found — so the reader can judge the strength of the identification instead of assuming it. Search the catalogue by CAS number to see one.
Note: InChI is developed by IUPAC and the InChI Trust. CAS Registry Number is a registered trademark of the American Chemical Society. MolGod.org is not affiliated with, endorsed by or acting on behalf of any of the organisations named on this page.
