Single-use words provide mathematical corrections to classical linguistic laws
Tracking the proportion of words that appear only once in a text allows researchers to adjust century-old scaling rules for vocabulary growth and rank distributions.

Across any novel, essay, or extended conversation, human language displays an unexpected statistical regularity. A small collection of grammatical words such as articles and conjunctions appears on nearly every page, while an enormous share of an author's total vocabulary surfaces only a single time throughout an entire work. Even in vast multi-volume libraries, thousands of distinct terms appear once and never return.
For over a century, quantitative linguists have sought mathematical formulas to describe how words distribute themselves across documents. Two classic formulations have long anchored the field. Zipf's law describes the relationship between the frequency of a word and its rank on an ordered frequency list, predicting that the second most common word appears roughly half as often as the first. Heaps' law describes how vocabulary size grows as a document lengthens, predicting that new words appear at a diminishing rate governed by a simple power curve.
While both classical formulas offer convenient approximations, real natural language texts routinely diverge from these clean power laws. In a preprint posted to the arXiv repository, researcher Łukasz Dębowski introduced systematic mathematical corrections to Zipf's law and Heaps' law by focusing on the proportion of once-occurring words.1
Why do single-use words dominate written documents?
Single-use words make up a substantial fraction of every natural language document because human authors draw from vast mental lexicons to express specific ideas. A word that appears exactly once in a given piece of writing is known to linguists as a hapax legomenon, often shortened to a hapax. When a reader opens a book, almost every initial word is new to that specific text. As reading continues, familiar words recur frequently, but the author continuously introduces rare descriptors, names, and specialized nouns.
Before calculating any scaling exponent, the process of vocabulary accumulation can be understood through a sequence of ordinary steps. First, an author begins writing by selecting words to communicate an initial concept. Next, grammatical requirements force structural words such as prepositions and pronouns to repeat immediately. Then, as the narrative progresses into new topics, the author must retrieve less common terms from memory. Finally, because specific descriptive words are rarely needed twice in the same context, a large pool of unique terms remains permanently tied to isolated occurrences. The rate at which these unique terms accumulate dictates the mathematical shape of the entire vocabulary.
The mathematical derivation in the preprint rests on two central assumptions.1 The first assumption relies on a standard urn model, a probability framework where words in shorter segments of text behave as if they were drawn blindly at random from a larger source document.1 The second assumption posits that the proportion of hapaxes follows a simple mathematical function of overall text length.1
By connecting the probability of drawing single-use words to the total length of a document, the framework links the local behavior of rare terms directly to the global rate of vocabulary accumulation across the whole text.

How do hapax rates adjust classical vocabulary scaling?
Tracking the proportion of single-use words provides explicit correction terms that refine the predictions of both Zipf's law and Heaps' law. Rather than assuming that word frequencies adhere strictly to an unbending power law across all ranks and lengths, Dębowski showed that tracking the hapax rate allows the classical curves to adjust dynamically as a text expands.
Dębowski evaluated four specific functional forms for the rate of once-occurring words: a constant model, a cancellation model, a linear model, and a logistic model.1 Each model represents a distinct hypothesis about how the proportion of single-use words shifts as the total word count increases from brief excerpts to full-length manuscripts.
The constant model assumes that the fraction of unique words never changes as a document grows. The cancellation model and linear model introduce progressive decreases in unique words as common terms accumulate. The logistic model describes an S-shaped transition where the rate of unique words begins with a steady plateau, drops as core vocabulary repeats, and gradually levels off toward a lower bound.
Dębowski explained the origin of these functional forms in response to questions from Primary. Dębowski said that his initial goal was to build on classical urn theory and understand why earlier corrections failed: “My initial goal was a bit modest. I wanted to develop further the classical theory of the urn model developed by Samuel Karlin, Estate Khmaladze and Harald Baayen. In particular, I wished to understand why the Zipf's law correction by Benoit Mandelbrot fails.”contributed
Dębowski added, “On my path, I discovered the hapax rate plot as a diagnostic tool. It showed so obviously why Mandelbrot's model fails, as it predicts a constant hapax rate. The other three parametric models were motivated mostly by the beauty of formulas, previous discussion in literature, or rough visual agreement with the empirical hapax rate plot.”contributed
What do mathematical models reveal about human vocabulary limits?
The logistic model produced the most accurate mathematical description of how single-use word rates behave across texts of varying lengths. Testing the four mathematical formulations against a sample of 14 texts written in English revealed that the logistic model yielded the best fit.1
The success of the logistic function points to cognitive constraints in individual writers. Dębowski said in response to questions from Primary that “the potential linguistic mechanism is plausibly connected to a finite working lexicon of a text author.”contributed In single-author works, a writer eventually exhausts their immediate working vocabulary, causing the arrival of genuinely novel terms to slow down predictably over the course of a manuscript.

Connecting the hapax trajectory to vocabulary limits addresses a foundational problem in language modeling. As Dębowski said, “In fact, the logarithm of the lexicon size is given as the area under the hapax rate function. If this area is finite then the lexicon is finite as well.”contributed Dębowski noted that models incorporating a bounded vocabulary match novel-length works with high precision.
Where does the single-author framework encounter its limits?
The mathematical derivations in the preprint describe idealized sampling conditions in individual books and do not represent a universal law for every collection of writing. The framework assumes that words can be modeled as if drawn blindly from an urn, a simplification that ignores how human discourse clusters words around evolving topics. Dębowski said that “the urn model usually gives a visually good prediction for novel-sized texts,” adding that larger discrepancies occasionally emerge when tracking vocabulary growth in initial text segments.contributed
The empirical evaluation in the preprint was conducted on a sample of 14 English texts, meaning the demonstrated fit reflects a specific small corpus. The preprint was posted to arXiv on August 21, 2026, and has not yet undergone formal peer review.1
The preprint also notes that larger corpora often exhibit two-regime vocabularies, where distinct frequency patterns govern common and specialized terms.1 Addressing those larger multi-author datasets requires more complex mixture models rather than a single simple equation.1
When moving from a single book to a massive archive composed by thousands of individuals, the statistical pattern shifts fundamentally. Dębowski said that “it is well known that the harmonic series is insummable. From this we can infer that Zipf's law of form f ~ 1/r must break for word ranks between 1,000 and 10,000.”contributed In massive multi-author collections, the hapax rate ceases to decline monotonically and instead forms a U-shaped curve, reflecting an open-ended cultural vocabulary that expands continuously across society.
Researchers can use the diagnostic shape of the hapax curve to decide when a dataset demands more sophisticated mixture models. As Dębowski noted, “Even a rough improvement of its modeling can translate to a large improvement in modeling Heaps' and Zipf's laws.”contributed Future investigations will test whether multi-regime mixture models can systematically unite the bounded lexicon of individual writers with the unbounded vocabulary of entire language communities.
Reporting note: This piece was prepared from the arXiv preprint and public records together with answers from Łukasz Dębowski to 7 questions from the Primary news team, completed August 2026.
References
This article is based on 1 source, with 7 statements from 1 contributor, listed in the order they are cited.
- 1 Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models See the source
- 2 Contributor commentary — Łukasz Dębowski 7 statements added to this article
Article history
-
7 statements 26 Aug 2026, 10:58What was added
The first assumption relies on a standard urn model, a probability framework where words in shorter segments of text behave as if they were drawn blindly at random from a larger source document.
On the record as reference 2What was addedIn a preprint posted to the arXiv repository, researcher Łukasz Dębowski introduced systematic mathematical corrections to Zipf's law and Heaps' law by focusing on the proportion of once-occurring words.
On the record as reference 2What was addedAddressing those larger multi-author datasets requires more complex mixture models rather than a single simple equation.
On the record as reference 2What was addedThe preprint also notes that larger corpora often exhibit two-regime vocabularies, where distinct frequency patterns govern common and specialized terms.
On the record as reference 2What was addedTesting the four mathematical formulations against a sample of 14 texts written in English revealed that the logistic model yielded the best fit.
On the record as reference 2What was addedThe second assumption posits that the proportion of hapaxes follows a simple mathematical function of overall text length.
On the record as reference 2What was addedDębowski evaluated four specific functional forms for the rate of once-occurring words: a constant model, a cancellation model, a linear model, and a logistic model.
On the record as reference 2ŁD Łukasz Dębowski · Contributor Researcher. -
Published 28 Aug 2026, 09:54Assembled by the Primary desk from 1 source · 16 cited sentences