Back to blogs
Optimization
Implementation
Coding
Concepts

What Is Tiktoken? Token Counting Explained (2026)

September 27, 2026
13 min read
What Is Tiktoken? Token Counting Explained (2026)
Share:

What Is Tiktoken? A Practical Guide to Token Counting, Encoding and LLM Costs in 2026

Large language models do not process text as humans do. Before a model can work with a prompt, the text is converted into tokens, and those tokens become the numerical representation the model processes. That makes token counting important for context limits, API costs, RAG systems, embeddings, agent workflows and prompt engineering.

Tiktoken is OpenAI's fast, open-source BPE tokenizer for OpenAI models. It converts text into token IDs and can convert those IDs back into text. It also provides model-aware encoding lookup through functions such as encoding_for_model(). The current PyPI release is 0.14.0, released August 17, 2026.

The important distinction is between counting a string and counting a complete API request. A tiktoken count tells you how a particular string is tokenized under a selected encoding. OpenAI's current documentation says that complete request counts can also include message structure, tools, schemas, images, files and conversations.

QUICK ANSWER

Tiktoken is OpenAI's fast byte pair encoding tokenizer for converting text into token IDs. You can install it with pip install tiktoken, choose an encoding such as o200k_base or cl100k_base, encode text, and count the resulting IDs with len().

Basic example:

import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
text = "Hello, how many tokens are here?"
tokens = encoding.encode(text)
print("Token count:", len(tokens))

Tokens are not words. A token can be a complete word, part of a word, punctuation or another text fragment. OpenAI gives a rough English estimate of about four characters or three-quarters of a word per token, but actual counts vary by text, language and encoding.

Use tiktoken for tokenizer-level estimates, choose the encoding that matches your target model, and use the API's official token-counting or usage information when you need the exact size or cost of a complete request.

1. What Is Tiktoken?

Tiktoken is an open-source tokenizer maintained by OpenAI. Its job is to transform text into tokens using byte pair encoding, or BPE. In a typical workflow, a string is encoded into integer token IDs, those IDs are supplied to the model, and decoding converts token IDs back into text.

It is designed for speed as well as correctness. OpenAI's repository reports that tiktoken is between three and six times faster than a comparable open-source tokenizer in the benchmark described in its README.

For developers, that makes tiktoken more than a token counter. It is a practical building block for prompt budgeting, document chunking, RAG, embedding preparation, context management and cost estimation.

2. Why Token Counting Matters

Token Count Matters Across AI Areas

OpenAI also notes that a cheaper price per million tokens does not automatically mean a cheaper completed task. Models can tokenize the same text differently and may generate different amounts of output or reasoning.

3. What Exactly Is a Token?

A token is a unit produced by a tokenizer. It is not a synonym for word, character or sentence.

Token Behavior by Input Type

Capitalization and surrounding spaces can also affect tokenization. OpenAI specifically notes that strings such as red, Red and a version with a leading space are not necessarily represented identically.

4. Tokens vs Words

A word counter is useful for editors, but it is not an exact LLM token counter. OpenAI's rough English estimate is one token for about four characters or about three-quarters of a word. That means 100 tokens is roughly 75 English words, but this is only a rule of thumb.

Technical writing, source code, JSON, tables, URLs and non-English text can have very different token-to-word ratios. If a program depends on a context limit or budget, count the actual text with the target tokenizer instead of multiplying the word count by a fixed number.

5. How BPE Tokenization Works

BPE, or byte pair encoding, represents text as reusable pieces. Instead of requiring every possible word to be stored as a single vocabulary item, the tokenizer can reuse common subword patterns.

OpenAI describes the tiktoken BPE design as reversible and lossless, capable of handling arbitrary text, and able to compress text so the resulting token sequence is shorter than the corresponding raw byte sequence. It also tends to expose recurring subword patterns such as common pieces inside words.

The exact pieces depend on the encoding. Therefore, the same sentence can produce different token counts under different tokenizers.

6. Tiktoken Encodings Explained

Encoding Models Reference Table

The OpenAI Cookbook documents the classic o200k_base, cl100k_base, p50k_base and r50k_base mappings. The current tiktoken model mapping also includes newer prefixes such as GPT-5, GPT-4.1, o-series and gpt-oss.

Do not select an encoding because its name looks familiar. Use encoding_for_model() when the model is recognized, or follow the target model's documentation.

7. Install Tiktoken

Install the package with:

pip install tiktoken

PyPI currently lists version 0.14.0 as the latest release, published August 17, 2026, with Python 3.9 or newer required.

For production applications, pin the package version in your dependency file so a future tokenizer-library update does not unexpectedly change application behavior.

8. Count Tokens in Python

The core operation is simple: load an encoding, encode the string, then count the resulting token IDs.

import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
text = "Tiktoken makes token counting easy."
tokens = encoding.encode(text)
print(tokens)
print("Token count:", len(tokens))

The returned list contains integer token IDs. The length of that list is the tokenizer-level token count for the supplied string.

9. Use encoding_for_model()

When tiktoken recognizes your model, encoding_for_model() is more convenient than manually selecting an encoding.

import tiktoken

model = "gpt-4o"
encoding = tiktoken.encoding_for_model(model)
text = "Count this prompt."
print(len(encoding.encode(text)))

The current model mapping includes prefixes for several modern model families, so a model-specific lookup can continue to work across dated model variants. If a model is unknown, tiktoken raises a KeyError and you should explicitly select the documented encoding.

10. Encode and Decode Tokens

Tiktoken supports both directions.

encoding = tiktoken.get_encoding("o200k_base")
text = "Tokenization is reversible."
tokens = encoding.encode(text)
restored = encoding.decode(tokens)
print(restored)

This is useful for token inspection, debugging prompt construction and understanding how a particular string is segmented.

11. Inspect Individual Token Pieces

To understand why a prompt uses a particular number of tokens, inspect the pieces themselves.

encoding = tiktoken.get_encoding("o200k_base")
text = "Hello, tokenization!"
for token_id in encoding.encode(text):
    print(token_id, repr(encoding.decode([token_id])))

This is especially helpful with code, punctuation-heavy prompts, identifiers, URLs and multilingual text.

12. Count Tokens in a Text File

from pathlib import Path
import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
text = Path("document.txt").read_text(encoding="utf-8")
count = len(encoding.encode(text))
print(f"Tokens: {count}")

Count the text after preprocessing. If your application extracts a PDF, removes HTML, normalizes whitespace, adds metadata or inserts separators, the final processed string is what matters.

13. Tiktoken for RAG

RAG systems are one of the clearest places where token counting becomes architecture. The model has to receive instructions, the user question, retrieved evidence, conversation history and enough output space to answer.

RAG Token Budget Components Table

A token-aware retriever can rank candidate chunks and add them until a safe context budget is reached. This is more robust than retrieving an arbitrary number of documents.

For the broader context-management approach, read What Is Context Engineering? Complete Guide (2026).

14. Tiktoken for Embeddings

OpenAI recommends tiktoken for estimating the token count of a string before sending it to an embedding model. Its current help documentation specifically notes cl100k_base for third-generation embedding models such as text-embedding-3-small and text-embedding-3-large.

This is useful in ingestion pipelines. If a document exceeds the model's input limit, split it before embedding. Token-aware chunking is more reliable than character-count chunking when your downstream model has token-based limits.

15. Token Counting and API Costs

Token counting helps estimate cost, but the price depends on the model and token category. OpenAI's current pricing separates input, cached input, cache writes and output for supported models. Some models also have different pricing for short and long context or processing modes.

Token Categories and Meanings Table

For a simple plain-text prompt, tiktoken gives a useful estimate. For final accounting, use the API usage data or official input-token counting mechanism for the endpoint.

16. Why Tiktoken Counts Can Differ From API Usage

Suppose you run len(encoding.encode(prompt)). That number represents the supplied string. It does not automatically represent the complete serialized API request.

Token Counting Methods Infographic

OpenAI explicitly notes that complete Responses inputs can include messages, images, files, tools and conversations, and that formatting tokens for request structure can also matter.

This is the single most important distinction to remember when building a token counter.

17. Does Tiktoken Count Images and Files?

Tiktoken is a text tokenizer. It does not turn a multimodal API request into one universal text-token number. Images and files have their own request accounting rules.

OpenAI's current documentation says its complete Responses input-token counting supports messages, images, files, tools and conversations. That is different from simply running tiktoken over a text string.

For multimodal applications, use the official request-level token-counting mechanism when you need the exact count.

18. Reasoning Tokens

Reasoning models can use internal tokens before producing the visible answer. Those tokens are not necessarily visible in the returned text, but OpenAI says they count toward output usage and are billed as output tokens.

This explains why a short visible answer can still consume more tokens than its displayed text suggests. Tiktoken cannot infer hidden reasoning-token usage from the final answer. The API's usage information is the correct source for that number.

19. Tiktoken for Prompt Engineering

Every instruction, example, schema and piece of conversation history consumes context.

Prompt Optimization Components Table

This connects directly to Model Routing for AI Coding Agents because token economics are part of routing decisions.

20. Tiktoken for AI Agents

Agents can generate much more token traffic than a single chatbot request. One task may trigger multiple reasoning calls, tool results, retries, summaries and follow-up calls.

Tiktoken for AI Agents

For agent architecture, see What Is an AI Agent? Beginner Guide With Examples (2026).

21. Common Tiktoken Mistakes

  • Treating words as tokens.
  • Using the wrong encoding for the model.
  • Counting only the prompt text and ignoring request structure.
  • Assuming tiktoken can calculate hidden reasoning usage.
  • Using character limits for token-constrained RAG chunks.
  • Calling a tokenizer count an exact API bill.
  • Hard-coding outdated model-to-encoding mappings.
  • Filling a context window just because the model supports it.
  • Comparing model prices without measuring total tokens generated.
  • Failing to count tool outputs and repeated agent calls.

22. Build a Reusable Token Counter

import tiktoken

def count_tokens(text: str, model: str = "gpt-4o") -> int:
    encoding = tiktoken.encoding_for_model(model)
    return len(encoding.encode(text))

print(count_tokens("Build a token-aware AI application."))

For a production implementation, handle unknown models explicitly and keep the model configurable. For exact request accounting, pair this utility with the endpoint's official token-counting or usage response.

23. Tiktoken vs Online Token Counters

Tiktoken vs Online Token Counters

For software systems, local programmatic counting is usually more useful because it can be integrated into validation, logging, routing and chunking.

24. Is Tiktoken Free and Open Source?

Yes. The tiktoken source repository is public on GitHub under the MIT license, and the package is distributed through PyPI.

Running the tokenizer locally does not itself create an OpenAI API charge. API costs occur when you make paid model requests or use other billable services.

25. Final Verdict: Is Tiktoken Worth Learning?

Yes. For developers working with OpenAI models and LLM applications, tiktoken is one of the most useful small libraries to understand.

It gives you a direct way to inspect tokenization, estimate prompt size, design token-aware chunking, protect context budgets and improve cost visibility. The current PyPI release is 0.14.0, released August 17, 2026.

The deeper lesson is that token counting is not just a billing trick. It affects context engineering, RAG quality, agent reliability, latency and model routing. At the same time, a tokenizer count is not the same as complete API usage. Structured messages, tools, files, images and reasoning can change the actual request usage.

My rating: 9.5/10 for developer usefulness, 9.5/10 for simplicity, 9/10 for performance and 9.4/10 overall.

Bottom line: learn tiktoken if you build LLM applications. Count the actual text you plan to send, use the correct encoding, budget context deliberately and rely on API usage information for final accounting.

Frequently Asked Questions

What is tiktoken?

Tiktoken is OpenAI's fast BPE tokenizer for converting text into token IDs and back into text.

How do I install tiktoken?

Run pip install tiktoken. PyPI currently lists version 0.14.0, which requires Python 3.9 or newer.

How do I count tokens with tiktoken?

Load an encoding, call encoding.encode(text), then use len() on the returned token list.

What is the difference between tokens and words?

A token can be a complete word, word fragment, punctuation or another text piece. Counts vary by encoding and language.

Which tiktoken encoding should I use?

Use encoding_for_model() when the model is recognized. Otherwise use the encoding documented for your model.

What is o200k_base?

It is a modern tiktoken encoding used by several newer OpenAI model families.

What is cl100k_base?

It is an encoding associated with GPT-4 and GPT-3.5-era models and several embedding models.

Can tiktoken calculate API cost?

It can estimate text-token usage, but complete request billing can include other content and different token categories.

Can tiktoken count reasoning tokens?

No. Hidden reasoning usage should be taken from API usage information.

Can tiktoken count images?

Tiktoken itself tokenizes text. Multimodal request accounting is handled by the model and API.

Is tiktoken free?

Yes. The open-source package is distributed under the MIT license.

Is tiktoken useful for RAG?

Yes. It helps control chunk sizes, retrieval budgets and context packing.

Is tiktoken useful for embeddings?

Yes. OpenAI recommends it for estimating embedding input size before processing.

What is the latest tiktoken version?

PyPI currently lists 0.14.0, released August 17, 2026.

Resources & Community

Join our community of 70,000+ AI enthusiasts and learn to build powerful AI applications. Build Fast with AI helps creators, developers and teams understand and implement practical AI.

Agentic AI Launchpad 2026

A structured 6-week cohort program that takes you from AI basics to building and deploying real-world agentic AI systems. Includes live sessions, expert mentorship, project reviews and a builder community network.

Ready to go from learning to building? Join the next cohort: Agentic AI Launchpad 2026

Free AI Resources

Access free tools, workshops and micro-learning to keep building.

References

Share: