PlumX Metrics
Embed PlumX Metrics

Text mixing shapes the anatomy of rank-frequency distributions

Physical Review E - Statistical, Nonlinear, and Soft Matter Physics, ISSN: 1550-2376, Vol: 91, Issue: 5, Page: 052811
2015
  • 23
    Citations
  • 96
    Usage
  • 17
    Captures
  • 1
    Mentions
  • 0
    Social Media
Metric Options:   Counts1 Year3 Year

Metrics Details

Most Recent News

Text mixing shapes the anatomy of rank-frequency distributions.

Authors: Jake Ryland Williams, James P Bagrow, Christopher M Danforth, Peter Sheridan Dodds PMID: 26066216 DOI: 10.1103/PhysRevE.91.052811 ISSN: 1550-2376 Journal Title: Physical review. E, Statistical,

Article Description

Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law, which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this "law" of ranks has been found to hold across disparate texts and forms of data, analyses of increasingly large corpora since the late 1990s have revealed the existence of two scaling regimes. These regimes have thus far been explained by a hypothesis suggesting a separability of languages into core and noncore lexica. Here we present and defend an alternative hypothesis that the two scaling regimes result from the act of aggregating texts. We observe that text mixing leads to an effective decay of word introduction, which we show provides accurate predictions of the location and severity of breaks in scaling. Upon examining large corpora from 10 languages in the Project Gutenberg eBooks collection, we find emphatic empirical support for the universality of our claim.

Provide Feedback

Have ideas for a new metric? Would you like to see something else here?Let us know