N-Gram Language Model Lab

    Learn about tokenization, build n-grams, and predict next words

    Runs in your browser. Nothing is uploaded.No signup requiredBuilt byDATAMETA LAB

    1Tokenization

    2N-Gram Builder

    3Tiny Predictor

    How to use this tool

    1. Enter your corpus text in Step 1
    2. Apply tokenization options (lowercase, remove punctuation)
    3. Click Tokenize to split text into tokens
    4. Choose N-gram type (Unigram, Bigram, or Trigram) in Step 2
    5. Build n-grams to see frequency table
    6. Enter a word/sequence in Step 3 to get predictions

    About this tool

    This interactive N-Gram Language Model Lab helps you understand the fundamentals of Natural Language Processing (NLP). N-grams are contiguous sequences of n items from a text. They're used in text prediction, spell checking, machine translation, and many other NLP tasks. Unigrams (N=1) are single words, bigrams (N=2) are pairs of consecutive words, and trigrams (N=3) are triplets. The lab demonstrates how language models use statistical patterns in text to predict the next word based on context.

    Frequently asked questions

    What is an n-gram?

    A run of n consecutive items from a text. In 'the cat sat', the unigrams are 'the', 'cat' and 'sat'; the bigrams are 'the cat' and 'cat sat'; the single trigram is 'the cat sat'. Counting them across a large corpus is the simplest way to model which words tend to follow which.

    How does an n-gram model predict the next word?

    It looks at the last n minus one words, finds every continuation that followed that sequence in the training text, and ranks them by how often each occurred. There is no understanding involved, only counting, which is exactly what makes the mechanism easy to see.

    What does tokenization do and why does it matter?

    It splits raw text into the units the model counts. The choices are consequential: lowercasing merges 'The' and 'the' into one token, which gives better statistics from a small corpus, and stripping punctuation stops 'cat.' and 'cat' being treated as different words. Every decision trades detail against data density.

    Why does a bigram model need more text than a unigram model?

    Because the number of possible sequences explodes with n. A vocabulary of a thousand words has a thousand unigrams but a million possible bigrams and a billion trigrams. Most never appear, so a higher n gives sharper predictions where you have data and no prediction at all where you do not. That trade-off is called sparsity.

    How do n-grams relate to modern language models?

    They are the direct ancestor. Both predict the next token from context; the difference is that an n-gram model can only look back a fixed number of words and can only recall exact sequences it has seen, while a neural model learns a representation that generalises to phrasings it has never encountered. Understanding the counting version makes the neural version far less mysterious.

    Is my text sent to a server?

    No. Tokenizing, counting and predicting all run in your browser. Whatever corpus you paste stays on your device.