N-grams from balanced National Corpus of Polish

The resource is a set of N-grams extracted from balanced National Corpus of Polish for N from 1 to 5. Each unigram is maximum continuous chunk of non-whitespace lower-case characters. The resource contains all unique N-grams followed by number of occurrencies.

NKJPNGrams

Menu

N-grams from balanced National Corpus of Polish

Download