N-Gram Language Model Lab
Learn about tokenization, build n-grams, and predict next words
1Tokenization
2N-Gram Builder
3Tiny Predictor
How to use this tool
- Enter your corpus text in Step 1
- Apply tokenization options (lowercase, remove punctuation)
- Click Tokenize to split text into tokens
- Choose N-gram type (Unigram, Bigram, or Trigram) in Step 2
- Build n-grams to see frequency table
- Enter a word/sequence in Step 3 to get predictions
About this tool
This interactive N-Gram Language Model Lab helps you understand the fundamentals of Natural Language Processing (NLP). N-grams are contiguous sequences of n items from a text. They're used in text prediction, spell checking, machine translation, and many other NLP tasks. Unigrams (N=1) are single words, bigrams (N=2) are pairs of consecutive words, and trigrams (N=3) are triplets. The lab demonstrates how language models use statistical patterns in text to predict the next word based on context.
Frequently asked questions
What is an n-gram?
A run of n consecutive items from a text. In 'the cat sat', the unigrams are 'the', 'cat' and 'sat'; the bigrams are 'the cat' and 'cat sat'; the single trigram is 'the cat sat'. Counting them across a large corpus is the simplest way to model which words tend to follow which.
How does an n-gram model predict the next word?
It looks at the last n minus one words, finds every continuation that followed that sequence in the training text, and ranks them by how often each occurred. There is no understanding involved, only counting, which is exactly what makes the mechanism easy to see.
What does tokenization do and why does it matter?
It splits raw text into the units the model counts. The choices are consequential: lowercasing merges 'The' and 'the' into one token, which gives better statistics from a small corpus, and stripping punctuation stops 'cat.' and 'cat' being treated as different words. Every decision trades detail against data density.
Why does a bigram model need more text than a unigram model?
Because the number of possible sequences explodes with n. A vocabulary of a thousand words has a thousand unigrams but a million possible bigrams and a billion trigrams. Most never appear, so a higher n gives sharper predictions where you have data and no prediction at all where you do not. That trade-off is called sparsity.
How do n-grams relate to modern language models?
They are the direct ancestor. Both predict the next token from context; the difference is that an n-gram model can only look back a fixed number of words and can only recall exact sequences it has seen, while a neural model learns a representation that generalises to phrasings it has never encountered. Understanding the counting version makes the neural version far less mysterious.
Is my text sent to a server?
No. Tokenizing, counting and predicting all run in your browser. Whatever corpus you paste stays on your device.

