[DSRP Evidence](https://dsrpevidence.org/)

# A comparative study on the names of higher education institutions in China and the United States: quantitative linguistic approaches

Zhu et al., 2026, Humanities and Social Sciences Communications — Linguistics

Patterns: [Systems](https://dsrpevidence.org/pattern/systems), [Relationships](https://dsrpevidence.org/pattern/relationships)

## In short

The whole corpus of names exhibits a statistical regularity invisible in any individual name, and a macro-level governance structure is treated as a relational driver of that diversity.

## What they found (results)

Corpus analysis of higher-education institution names in China and the US found both national naming corpora follow a Zipf-Mandelbrot frequency distribution, with the US corpus showing higher naming-diversity entropy than China's, a difference the authors link to more decentralized, market-oriented university governance in the US.

## Abstract

This study adopts quantitative methods to analyze the linguistic features of the names of higher education institutions in China and the United States (US). The findings reveal both similarities and differences in the university names of the two countries. Firstly, the Zipf-Mandelbrot law is found to govern the distributions of university names in both countries. Secondly, in terms of word frequency, the findings show that US university names contain more personal names and religious terms than Chinese university names. In addition, the results of entropy values show that the naming pattern of US universities is more diversified than that of Chinese universities, which may be related to the decentralized, market-oriented governance structure of US higher education and its multicultural social context. We discuss the practical implications for natural language processing tasks, such as named-entity recognition and vocabulary compression. We also discuss the methodological implications for future onomastic research, highlighting the value of integrating corpus linguistic techniques with quantitative onomastics.

These researchers were not testing DSRP. The finding is theirs; the correspondence to DSRP is drawn by this site.

[Source](https://doi.org/10.1057/s41599-026-08701-y)
