## Statistical properties ### Heap's law - Estimate the vocabulary size of a collection - `$ T $`: Number of tokens - `$ M $`: Vocabulary size (distinct terms in collection) - `$ M=kT^b $` - `$ 30\leq k \leq 100 $` - `$ b\approx 0.5 $` ### Zipf's law - Distribution of terms in a collection - Collection frequency inverse proportional to rank in frequency table - Sort terms by frequency `$ t_1,..., t_M $` - `$ cf(t_i) \propto ci^{-1}, c\in \mathbb{R} $` - `$ cf(t_i) = c * i^{-1} $` - `$ \log cf(t_i) = \log c - \log i $` - Example - Second most frequent word has half of the occurrences of the first - Third most frequent word has third of the occurrences of the first - Many types of data in physical and social science approximated by Zipf's law