Verification of the Heaps law using the Google Books Ngram database
Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki, Uchenye Zapiski Kazanskogo Universiteta. Seriya Fiziko-Matematicheskie Nauki, Tome 155 (2013) no. 4, pp. 16-23

Voir la notice du chapitre de livre provenant de la source Math-Net.Ru

This article is devoted to the verification of the Heaps empirical law for European languages using the Google Books Ngram corpus data. It is shown that the Heaps law holds only for short texts and texts related to short historical periods. The Heaps exponent decreases in time and varies significantly within characteristic intervals of 60–100 years. The relationship between the word frequency distribution and the expected dependence of the number of individual words on the text size is analyzed in terms of a simple probability model of text generation. This model serves as an explanation for the observed decreasing trend of the Heaps exponent.
Keywords: Heaps law, Zipf law, text probability models, Google Books Ngram corpus.
@article{UZKU_2013_155_4_a1,
     author = {V. V. Bochkarev and E. Yu. Lerner and A. V. Shevlyakova},
     title = {Verification of the {Heaps} law using the {Google} {Books} {Ngram} database},
     journal = {U\v{c}\"enye zapiski Kazanskogo universiteta. Seri\^a Fiziko-matemati\v{c}eskie nauki},
     pages = {16--23},
     publisher = {mathdoc},
     volume = {155},
     number = {4},
     year = {2013},
     language = {ru},
     url = {http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/}
}
TY  - JOUR
AU  - V. V. Bochkarev
AU  - E. Yu. Lerner
AU  - A. V. Shevlyakova
TI  - Verification of the Heaps law using the Google Books Ngram database
JO  - Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki
PY  - 2013
SP  - 16
EP  - 23
VL  - 155
IS  - 4
PB  - mathdoc
UR  - http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/
LA  - ru
ID  - UZKU_2013_155_4_a1
ER  - 
%0 Journal Article
%A V. V. Bochkarev
%A E. Yu. Lerner
%A A. V. Shevlyakova
%T Verification of the Heaps law using the Google Books Ngram database
%J Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki
%D 2013
%P 16-23
%V 155
%N 4
%I mathdoc
%U http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/
%G ru
%F UZKU_2013_155_4_a1
V. V. Bochkarev; E. Yu. Lerner; A. V. Shevlyakova. Verification of the Heaps law using the Google Books Ngram database. Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki, Uchenye Zapiski Kazanskogo Universiteta. Seriya Fiziko-Matematicheskie Nauki, Tome 155 (2013) no. 4, pp. 16-23. http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/