BlogA unit of exchange n...

A unit of exchange nobody defined. We need to standardize the AI token.

A unit of exchange nobody defined. We need to standardize the AI token.

I need you to take a moment and think about the second. It is the SI unit for time.

The International System of Units defines it with exact precision: the duration of exactly 9,192,631,770 periods of the radiation corresponding to the transition between the two hyperfine levels of the ground state of the caesium-133 atom. Not approximately, not roughly, exactly! Every laboratory in the world, every telecommunications system, every financial settlement network, every GPS satellite agrees on what a second is. That agreement did not happen by accident. It happened because civilisation decided that a unit underpinning so much commerce and science was too important to leave undefined.

Now consider the token.

The token is the fundamental unit of the AI economy. It determines what you pay when you use an AI service. It determines how much information a model can hold in memory at once. It shapes how developers build applications, how researchers design experiments, how businesses forecast costs, and how governments procure AI infrastructure. Billions of dollars of contracts are being written in tokens. Entire technical architectures are being designed around tokens. The AI industry has staked its commercial model on the token.

Nobody has defined it.

Not in any way that holds across vendors or survives a model update. The most important unit in one of the most consequential industries in the world is, at this moment, undefined.

That is the problem I'm here to bitch about.

What a Token Actually Is

To understand why the lack of a standard matters, you first need to understand what a token actually is in technical practice, and why the answer is less satisfying than it should be.

Modern AI language models do not process text the way humans read it, word by word or letter by letter. They operate on chunks of text called tokens, produced by a preprocessing algorithm called a tokenizer. The tokenizer takes raw text and segments it into sub-word units, fragments that may be whole words, parts of words, punctuation marks, or individual characters, depending on context.

The most widely used family of tokenization algorithms is called Byte Pair Encoding, or BPE. It works by analysing a large corpus of text, identifying the most frequently co-occurring character sequences, and progressively merging them into single vocabulary entries. The resulting vocabulary (typically between 32,000 and 100,000 entries) becomes the model's alphabet. Every piece of text it will ever see must be expressed in terms of that vocabulary.

The rough rule of thumb you will encounter across the industry is that one token equals approximately four characters of English text, or about three-quarters of a word. OpenAI, Anthropic, and Google all land somewhere near this average for standard English prose. But that similarity is coincidence, not coordination. Each lab trained its tokenizer independently, on different corpora, with different vocabulary sizes and different algorithmic choices. The word "unbelievable" might be one token in one system and three in another. A space before a word can change how it tokenizes. Emojis, code syntax, and non-Latin scripts all behave differently across systems.

There is no shared definition. There is no reference standard. There is no conversion factor published by any of the major labs. There is just a loosely similar engineering approach producing meaningfully different outputs, all marketed under the same name.

The Economy Built on an Undefined Unit

This would be a minor technical curiosity if tokens were only used internally. They are not. Tokens are the unit of commercial exchange in the AI industry, and the consequences of their non-standardisation are concrete and growing.

Pricing opacity.

When Amazon charges you per kilowatt-hour, a kilowatt-hour is a kilowatt-hour. When a cloud provider charges per gigabyte of storage, a gigabyte is a gigabyte. When OpenAI charges per million tokens and Anthropic charges per million tokens, those are not the same million. A business trying to forecast AI infrastructure costs across providers is comparing currencies without an exchange rate, and nobody is publishing one. Cost modelling exercises that treat vendor tokens as equivalent are with all due respect, wrong.

Benchmarks without a baseline.

Context window comparisons are meaningless without a common unit. When one model advertises a one-million-token context window and another advertises two hundred thousand tokens, those numbers cannot be directly compared. The model with the larger window may actually process less real-world information if its tokenizer is less efficient. Cross-vendor performance benchmarks are contaminated by the same problem. The industry talks about tokens as if they are interchangeable. They are not.

The moving target problem.

Tokenizers change between model versions. The tokenizer underpinning GPT-3 was probably not the same as GPT-4's. Systems calibrated against one version may behave differently after a model update. Not because the model's reasoning changed, but because the unit of measurement shifted underneath the application. There is no versioning standard, no changelog obligation, no notification requirement. Developers absorb this instability silently.

Architectural lock-in.

Any developer building a RAG pipeline, a summarisation system, or an agentic workflow has to make token assumptions somewhere. In how they chunk documents, manage context, and estimate costs. Those assumptions are vendor-specific and version-specific. Switching providers is not just a business negotiation; it is a re-engineering project, partly because the fundamental unit does not transfer. This is not neutral market competition. It is incompatibility masquerading as differentiation.

The Language Inequality Problem

Of all the consequences of token non-standardisation, one deserves to be named loudly and separately, because it is the one most likely to be treated as a footnote when it should be a headline.

Tokenizer design is not a neutral technical decision. It is a decision about whose language is treated as the default, whose vocabulary gets compressed efficiently, and whose community pays a structural premium for access to AI.

English was the dominant language in most tokenizer training corpora. The result is that English text tokenizes efficiently. Roughly one token per four characters. A model with a two-hundred-thousand-token context window can hold a substantial amount of English text in memory. The same model, processing an equivalent amount of information expressed in Yoruba, Igbo, Hausa, or Swahili, may consume two to three times as many tokens to convey the same semantic content. The context window effectively shrinks. The cost per idea expressed rises. The API bill for the same cognitive task is higher.

This is not an edge case. It is not a minor inefficiency affecting small communities on the periphery of the AI market. Africa has over two thousand languages and a rapidly growing technology sector. Nigerian startups are building AI-powered products. Kenyan developers are integrating language models into financial services. African governments are beginning to procure AI infrastructure at scale. Every one of them is doing so from a structurally disadvantaged position. Not because of application-layer decisions or pricing tiers, but because of tokenizer design choices made by engineers in San Francisco and London who were not optimising for Igbo vocabulary.

This is not intentional discrimination. It is the emergent consequence of building a global standard on a narrow corpus. But the effect is the same. And unlike many inequities in the technology industry, this one is relatively tractable. It is a solvable engineering and policy problem. It just requires the communities most affected to be in the room when it is discussed. Currently, they are not.

Why This Has Not Been Solved

It is worth being honest about why this problem persists rather than simply noting that it does.

The commercial disincentive is real. Proprietary tokenizers create switching costs, and switching costs are moats. A developer who has built their application around one vendor's token assumptions has an implicit reason to stay. Standardisation would erode that lock-in, which is precisely why no major lab has volunteered to lead the effort.

The technical complexity is also real. Agreeing on a single tokenization standard would require negotiating vocabulary size, segmentation algorithm, handling of special characters, multilingual coverage minimums, and update protocols. It would require resolving genuine tradeoffs. A tokenizer optimised for multilingual equity may be slightly less efficient for English, and vice versa. These are tractable problems, but they are not trivial ones.

And standards bodies move slowly. The ITU took decades to standardise telecommunications protocols. The internet's foundational standards emerged through years of rough consensus and running code. The transformer-based language model era is less than a decade old in practical deployment terms. The argument that formal standardisation is premature is not entirely without merit.

But none of these obstacles is an argument for waiting. They are arguments for starting. The window in which the standard can be shaped (before it hardens around assumptions that were never designed with non-English communities in mind) is open now. It will not stay open indefinitely.

What Standardisation Could Look Like

Standardisation does not require every lab to use the same tokenizer. It requires common measurement and mandatory disclosure. Those are different, and considerably more achievable, demands.

A reference tokenizer.

A neutral standards body (NIST in the United States, ISO/IEC JTC 1 internationally) could define and maintain an open-source reference tokenizer: versioned, frozen at each release, publicly auditable. Vendors would not be required to use it in their models. They would be required to publish the conversion factor between their tokenizer and the reference, across a defined set of languages. This is analogous to how industrial measurement standards work. You do not require all manufacturers to use the same ruler. You require them to calibrate against a common standard and disclose the result.

Disclosure mandates.

Even before a reference tokenizer exists, regulators could require vendors to publish token efficiency ratios across languages relative to a declared baseline. The European Union's AI Act, already the most substantive AI regulatory framework in force, is a plausible vehicle for this kind of disclosure requirement. The UK's AI Safety Institute and the United States' NIST AI Risk Management Framework are others. These mechanisms exist. They need to be pointed at this problem.

Procurement leverage.

African governments and large enterprise buyers do not need to wait for international standards bodies. Procurement standards for AI services can require vendors to disclose token efficiency ratios across relevant languages as a condition of contract. If the Nigerian government, the Kenyan government, and the South African government all made this a procurement requirement, the major AI vendors would publish the data within a quarter. Procurement is a form of regulation that does not require legislative consensus.

The research bridge.

While formal standardisation lags, the research community can publish systematic cross-tokenizer equivalence tables — empirical exchange rates across major models for different languages and content types. These would not be standards, but they would be working tools. African AI research institutions are well-positioned to lead this work, and doing so would simultaneously raise the profile of the language equity dimension in international standards conversations.

The Bigger Argument

The token is not a technical detail. It is the unit of account for an economic system that is rapidly becoming critical infrastructure. For businesses, governments, researchers, and individuals worldwide.

We have never allowed critical infrastructure to self-define its own units of measurement without oversight. We do not let electricity providers define what a kilowatt-hour means. We do not let pharmaceutical companies define their own dosage units. We do not let financial institutions invent their own currencies and charge for services denominated in them without disclosure. The principle that commercial units of exchange must be standardised, transparent, and auditable is not controversial. It is foundational to how functional markets work.

The AI industry is not exempt from this principle. It is simply young enough, and has moved fast enough, that the institutional response has not caught up. That gap will close. The question is whether it closes with the right people at the table.

The communities with the most to lose from a bad standard being locked in are the ones building in Yoruba, Igbo, Amharic, and Hausa. They are the ones whose context windows are already effectively smaller, whose API costs are already structurally higher, and whose developers are already building on a foundation that was not designed with them in mind. They are also the ones currently least represented in the forums where AI infrastructure standards are being shaped.

The second was defined because precision mattered. The token will be defined eventually. By markets, by regulation, or by formal consensus. The only question is whether the communities most affected by that definition will have a seat at the table when it happens.

That seat will not be offered. It will have to be claimed.

Date Published: 2026-07-12

Enjoyed the article? Share