Verification of the Heaps law using the Google Books Ngram database
    
    
  
  
  
      
      
      
        
Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki, Uchenye Zapiski Kazanskogo Universiteta. Seriya Fiziko-Matematicheskie Nauki, Tome 155 (2013) no. 4, pp. 16-23
    
  
  
  
  
  
    
      
      
        
      
      
      
    Voir la notice du chapitre de livre provenant de la source Math-Net.Ru
            
              This article is devoted to the verification of the Heaps empirical law for European languages using the Google Books Ngram corpus data. It is shown that the Heaps law holds only for short texts and texts related to short historical periods. The Heaps exponent decreases in time and varies significantly within characteristic intervals of 60–100 years. The relationship between the word frequency distribution and the expected dependence of the number of individual words on the text size is analyzed in terms of a simple probability model of text generation. This model serves as an explanation for the observed decreasing trend of the Heaps exponent.
            
            
            
          
        
      
                  
                    
                    
                    
                    
                    
                      
Keywords: 
Heaps law, Zipf law, text probability models, Google Books Ngram corpus.
                    
                  
                
                
                @article{UZKU_2013_155_4_a1,
     author = {V. V. Bochkarev and E. Yu. Lerner and A. V. Shevlyakova},
     title = {Verification of the {Heaps} law using the {Google} {Books} {Ngram} database},
     journal = {U\v{c}\"enye zapiski Kazanskogo universiteta. Seri\^a Fiziko-matemati\v{c}eskie nauki},
     pages = {16--23},
     publisher = {mathdoc},
     volume = {155},
     number = {4},
     year = {2013},
     language = {ru},
     url = {http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/}
}
                      
                      
                    TY - JOUR AU - V. V. Bochkarev AU - E. Yu. Lerner AU - A. V. Shevlyakova TI - Verification of the Heaps law using the Google Books Ngram database JO - Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki PY - 2013 SP - 16 EP - 23 VL - 155 IS - 4 PB - mathdoc UR - http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/ LA - ru ID - UZKU_2013_155_4_a1 ER -
%0 Journal Article %A V. V. Bochkarev %A E. Yu. Lerner %A A. V. Shevlyakova %T Verification of the Heaps law using the Google Books Ngram database %J Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki %D 2013 %P 16-23 %V 155 %N 4 %I mathdoc %U http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/ %G ru %F UZKU_2013_155_4_a1
V. V. Bochkarev; E. Yu. Lerner; A. V. Shevlyakova. Verification of the Heaps law using the Google Books Ngram database. Učënye zapiski Kazanskogo universiteta. Seriâ Fiziko-matematičeskie nauki, Uchenye Zapiski Kazanskogo Universiteta. Seriya Fiziko-Matematicheskie Nauki, Tome 155 (2013) no. 4, pp. 16-23. http://geodesic.mathdoc.fr/item/UZKU_2013_155_4_a1/