36 | Legacies of the 1931 Linguistic Census
- Mar 14
- 15 min read
By Gaurav Kalyani and Shivakumar Jolad
Published on: 31 August 2026
Introduction
The 1931 census would prove to be the last full, undivided linguistic census of the Indian subcontinent. Partition in 1947 broke the statistical unit apart forever, and no later Indian census has attempted quite the same exercise of mapping mother-tongue and subsidiary-language distribution together at an all-India scale.
This essay surveys how language was enumerated and studied in the 1931 census, and traces the social and political consequences of that exercise, both in its own time and in the decades that followed.
Understanding Language Returns
The 1931 census leaned heavily on Sir George Grierson's monumental Linguistic Survey of India, whose many volumes including the introductory volume had appeared as recently as 1927. On that, Hutton commented, “little remains for a census to do with regard to the main languages of the country beyond recording their corresponding increase and decrease.”
A regular decennial census could offer an ongoing, changing picture of language use over time, which a one-time survey couldn’t provide. Most importantly, the census documented bilingualism between neighboring language groups, which was an area the initial survey largely left undocumented.
Two columns were provided on the household schedule. One was for mother-tongues and another for any subsidiary language used in daily life. Infants and the deaf and dumb were credited with the mother-tongue of their mothers.
This “subsidiary language” column was new in most provinces in 1931, and it is the reason the census could, for the first time, attempt to measure bilingualism systematically rather than forcing every respondent into a single linguistic box.
Hutton emphasized that documenting overlapping languages carried significant practical importance beyond mere academic interest. This data was highly relevant due to the reasonable desire among many Indians to reorganize provincial boundaries on a linguistic basis.
Commenting on this matter, Hutton states,
“In one respect, however, existing information was lacking and that was the extent of the overlap. of different languages in the numberless areas in which two or more co-exist. It is not suggested that this overlap is a permanent feature, and that areas speaking two languages at present will necessarily continue to do so in perpetuity, but in the view of the not unreasonable desire of many Indians for a redistribution of provinces on a linguistic basis, as well as of the possibility of extensions of franchise to very considerable populations speaking some tribal language as their mother-tongue, to say nothing of the desirability of starting all primary :ducation in the real language of the child to be taught, the record of this overlap has more than a purely academic interest.”
In other words, the census framers knew before a single schedule was filled that the numbers they gathered would be read as evidence in live political arguments about how India's map should be redrawn and who should get to vote.
Upon classification, the 1931 report counted 225 languages in India, three more than the 222 recorded in 1921. The net change in language classifications was driven by several additions, which were offset by the consolidation of Tibeto-Burman nomenclature. These adjustments included recording Persian in Baluchistan, recognizing Bashgali of the Kafir group, and reclassifying Konkani as its own language rather than a Marathi dialect. Additionally, Burushaski appeared for the first time, Andamanese was split into two distinct languages, and Gypsy dialects were divided into six separate categories.
Even this seemingly technical exercise in taxonomy was itself a form of political classification. Deciding whether to classify Konkani, Halbi, and Gondi as independent languages or dialects had major political consequences. These decisions directly affected how provincial borders were drawn and how communities were represented.
Table: Major language families of India by number of speakers, 1931
Family | Persons | % of population |
Indo-European (chiefly Indo-Aryan) | 257,492,805 | 73.5% |
Dravidian | 71,644,787 | 20.4% |
Tibeto-Chinese | 14,010,496 | 4.0% |
Austric (Munda, Mon-Khmer, Austronesian) | 5,342,708 | 1.5% |
Karen | 1,341,331 | 0.4% |
Unclassed and other | 302,324 | 0.1% |
Source: Census of India 1931, Table XV, Subsidiary Table I (undivided India, including princely states; excludes Vernaculars of foreign countries and Europe).
Statistics of Language in 1931 Census
The 1931 census report provides a structured statistical overview of language distribution across British India and Burma through its Subsidiary Table I. Out of approximately 366.4 million recorded speakers, which counted bilingual individuals twice, the vast majority of the population, i.e. around 261.1 million, spoke an Indo-Aryan language. The remaining population was divided among other major linguistic families, including over 79 million Dravidian speakers, 14.1 million Tibeto-Burman speakers along the northern and eastern frontiers, and 4.7 million Munda (Austroasiatic) speakers.
Overall, just nineteen Indo-Aryan languages accounted for nearly three-quarters of the subcontinent's total population, representing a concentrated linguistic footprint that saw an increase of over 50.9 million mother-tongue speakers across all language families since 1921.
At the individual level, Western Hindi remained the most widely spoken mother tongue with 71.5 million speakers, comprising approximately 37.7 million males and 33.8 million females. While this figure was lower than the 96.7 million recorded in 1921, the decline was merely administrative, resulting from the reclassification of Eastern Hindi and other intermediate linguistic groups rather than an actual drop in speakers.
Other prominent languages included Bengali with 53.4 million speakers, Bihari with 27.9 million, and Telugu with 26.3 million. Rounding out the ten most spoken languages in the subcontinent were Marathi (20.8 million), Tamil (20.4 million), Punjabi (15.8 million), Rajasthani (13.8 million), Kanarese (11.2 million), and Oriya (11.1 million).

Script
Language enumeration in 1931 could not avoid the question of script, and script could not avoid politics. Since the nineteenth century, the United Provinces has been at the center of an ongoing dispute over the region's spoken language- Hindi or Urdu. The debate was over whether this local tongue should be written, taught, and used in courts as "Hindi" in the Nagari script or "Urdu" in the Persian script.
Hutton records that
“in point of practice it is impossible to define any boundary between Urdu and Hindi as spoken, since the difference consists merely in a preference for a Persian or for a Sanskrit vocabulary.”
Rather than reopen the controversy, the 1931 schedule for the United Provinces instructed enumerators simply to record everyone as speaking “Hindustani”. This decision, as Hutton observed, “caused some searching of heart among Muslims” who feared the disappearance of the Urdu label from the record.
This was not a new fight, nor did recording “Hindustani” resolve it. Asha Sarangi (2009), in a detailed study of enumerative practice in the United Provinces, traces how census categories hardened the Hindi–Urdu distinction over successive censuses. In 1872 and 1881 the two were recorded interchangeably.
The 1901 census, in the wake of the April 1900 Nagari Resolution that had made Hindi a court language alongside Urdu, instructed enumerators for the first time to record Urdu separately and all other languages and dialects…as Hindi. This formula folded Kaithi, Braj and Khari Boli speakers into the Hindi column and produced a linguistic hegemony of Hindi over Urdu, as Sarangi asserted.
From 1901 onward, she argues, “the politics of numbers starts affecting the domain of cultural and political struggle over linguistic identities,” with script, mannerism and vocabulary increasingly read as markers of religious community such as, Nagari with Hindu, Persian with Muslim.
Gandhi's own intervention in the 1920s was to promote “Hindustani,” a fused, script-neutral vernacular, as a way of de-communalising the dispute. The choice of “Hindustani” as the recorded category in the 1931 United Provinces schedule was, in a sense, an administrative echo of that campaign (Sarangi, 2009).
But as Sarangi shows, the label did not dissolve the underlying rivalry. Hindi and Urdu partisans on both sides continued to treat Hindustani as a battleground to be annexed rather than a synthesis to be embraced.
The argument resurfaced, as Sarangi noted, with even greater bitterness in 1941, when the United Provinces government's failure to supply Urdu-language schedules provoked formal protest from Muslim leaders that Muslim respondents' mother tongue was being wrongly recorded as Hindi or Hindustani.
Beyond the Hindi–Urdu debate, script surfaced elsewhere too. Only a handful of provinces like Punjab, the Central Provinces and Berar, the Central India Agency, Hyderabad, and Jammu and Kashmir had collected any return of the script in which people were literate. It did not distinguish Nagari, Urdu (Persian) and Gurmukhi readers.
Hutton observed that “the need for a common script for India is probably even greater than that for a common tongue,” and floated the possibility that Hindustani in Roman script, which was already used with some success in the Indian Army, might eventually serve that purpose. This suggestion went nowhere but it reflected how open and contested the question of a national linguistic standard still was on the eve of the 1930s constitutional reforms.
Hutton also recorded a Hindi-nationalist newspaper's own slogan which was published in English — “Linguistic Inqilab Zindabad,” which he translated as “Up the Zabani Rebels.”
Linguistic Pre-History
In a section of the Language chapter of the 1931 census report, Hutton stepped away from demographic counting to explore the fields of philology and archaeology. A central focus of this section was the ancestry of the Brahmi alphabet, which serves as the parent script of almost all indigenous Indian writing.
The section begins by revisiting an 1867 hypothesis by E. Thomas, who argued that the Sanskrit alphabet was not a post-Aryan import but was instead derived from a writing system developed by India's pre-Aryan inhabitants. However, Hutton, based on the contemporary archaeological evidence, tried to establish a connection between the historical Brahmi alphabet and the symbols found on the pre-Aryan Mohenjodaro and Indus Valley seals. It is important to note that current scholarly consensus does not find any definitive connection between the two.
To determine what languages were spoken in Upper India before the Indo-Europeans arrived, Hutton evaluated the oldest language layers. Utilizing the principle that languages with the widest geographical distribution are typically the oldest in time, he concluded that the Austroasiatic (Munda) family represents an older group of tongues than Dravidian. Debates persist regarding whether the Munda languages originally spread into India from the east or the west.
Hutton notes that this linguistic mystery is closely tied to the prehistoric distribution of physical artifacts, specifically the "shouldered celt" (a distinctively shaped stone and metal adze or hoe).



Some scholars suggested these tools represent an oceanic intrusion from Indonesia and Polynesia. Others argue that the stone adzes were actually crude imitations of copper originals manufactured in the ancient, non-Aryan Munda kingdoms of India, which then spread eastward into the Pacific. Jean Przyluski identified numerous Austroasiatic loan words in modern Indo-European vocabularies and connected Munda populations to widespread regional myths, such as the gourd-origin folklore.
However, rejecting these theories, Hutton positions Dravidian speakers as the latest pre-Indo-European occupants of Upper India. He suggested that the Dravidians reached India from the north-west, bringing with them an advanced, urban civilization historically connected to Mesopotamia, Asia Minor, and the eastern Mediterranean. He supported his argument by connecting Brahui of Baluchistan and 1930 linguistic investigations showing striking structural similarities between Dravidian and Kharian (Hurrian), an ancient language of the Euphrates valley.
This section concluded by analyzing the distribution of Indo-European languages, famously split by Sir George Grierson into "Inner" and "Outer" bands. Rather than relying on theories of multiple, distinct waves of northern invasions, Hutton proposed that the "Outer Band" (including languages like Marathi, Bengali, and Sindhi) retain features of the ancient Dardic or Pisacha group.
He further posits that this linguistic footprint may have been left by Pamir Alpines who occupied the Indus valley after the fall of Mohenjodaro but before the Rigvedic Aryans arrived. When the Rigvedic Aryans ultimately colonized the Ganges basin, their language fused with the highly developed pre-existing cultures. It was this creative amalgamation of populations, rather than isolated martial conquest, that allowed Indo-European languages to achieve their greatest heights of written literature in India.
Language and Boundary Disputes: The Case of Orissa
The demand for a separate Oriya-speaking province that was eventually realised as Orissa on 1 April 1936, was India's first linguistic province. It relied substantially on census language figures, and the Orissa Boundary Committee requisitioned the raw schedule material from the Madras Superintendent of Census Operations before it had even been tabulated, forcing his staff to work, in Hutton's words, “to the verge of collapse” to meet the Committee's deadline.
Hutton noted that the accuracy of the census was compromised by misleading propaganda campaigns. These efforts deliberately misclassified individuals as either speaking or not speaking the Oriya language, thereby distorting the very linguistic data the census was designed to collect.
Bilingual Oriya–Telugu speakers, rather than honestly returning both languages in the mother-tongue and subsidiary columns, frequently “plumped” for a single language out of political loyalty. An Oriya-mother-tongue Telugu speaker concealed his Telugu, and vice versa, each side fearing that an honest bilingual return would be read as ceding ground to the rival community.
A Madras government enquiry into the equally contentious 1901 figures had already found that Telugu speakers in the area had incentives to return themselves as Oriya to access benefits reserved for Oriya education, while a numerical preponderance of Oriya-speaking enumerators tended to record ambiguous answers as Oriya by default.
A similar but smaller dispute occurred in Bhopal State, where the local administration intentionally ignored the facts during the census. Instead of recording the residents' actual Rajasthani or Gondi mother tongues, the administration registered both Hindu and Muslim subjects as Urdu speakers to align them with the state's preferred version of Hindustani.
Among Gond communities elsewhere, the reverse dynamic was observed. Some Gonds, regarding Gondi as socially inferior, returned Halbi (a Marathi dialect) as their mother-tongue even while their households continued to speak Gondi at home. Mr Grigson, an administrator of Bastar State, observed this pattern directly in court testimony, where husbands claiming to know “no language but Halbi” were contradicted by wives and mothers who understood only Gondi.
Tribal Languages: Documentation and Disappearance
The language chapter also devoted substantial space to the fate of India's tribal languages, a subject Hutton treated with particular interest. He was an anthropologist by training who had served extensively among Naga tribes in Assom. He argued that how well a tribal language survives is a better measure of a tribe's social unity than the survival of its religion. This is because he thought tracking language was easier and more definite than tracking religion, which often has vague and unclear boundaries with Hinduism.
In the United Provinces, gypsy languages were reported as dying out entirely, their speakers absorbed into settled agricultural life and Hindustani. In Central India, the Census Superintendent reported that Munda-speaking groups such as the Kol, Baiga and Saharia had “no languages of their own” left at all.
Yet in Bihar and Orissa and in Assam, tribal languages were reported as remarkably resilient. The Bihar and Orissa Superintendent recorded a 17.7 per cent increase in tribal-language speakers since 1921, attributing it partly to natural population growth and partly to the new subsidiary-language column itself, which allowed bilingual tribal respondents to record their mother-tongue honestly for the first time.
In one instance from Champaran district, not a single Oraon had been recorded as speaking Oraon in 1921 out of a population of nearly 10,000. However in 1931, with the two-column system, 5,511 were recorded as Oraon mother-tongue speakers and the great majority of them also fluently bilingual in Hindustani.
The Central Provinces data on bilingualism among tribal-language speakers offers a rare, granular statistical window onto this process.
The census showed that three out of every four persons in that province who spoke a subsidiary language in daily life belonged to one of eighteen listed tribes, with rates of bilingualism varying enormously by community.
Table: Bilingualism among selected tribal-language speakers, Central Provinces, 1931
Tribal language | Number of speakers | Per mille (1000) speaking a subsidiary language | percentage |
Birjia | 628 | 998 | 99.8 |
Korwa | 12,431 | 950 | 95 |
Kondh | 133,682 | ~880 | ~88 |
Kharia | 113,680 | 871 | 87.1 |
Asuri | 2,769 | 842 | 84.2 |
Gondi | 6,270 | 769 | 76.9 |
Mundari | 421,811 | 486 | 48.6 |
Ho | 526,443 | 226 | 22.6 |
Source: Hutton, 1933
Hutton's contributors also pointed out that the census both measured and, in a slower structural sense, participated in tribal-language decline.
The standard process of language shift occurs when a tribal mother tongue continues to be used at home for a generation or two while speakers adopt a regional language for trade, legal matters, and employment. As highlighted in several provincial reports, this transitional bilingualism serves as the first stage toward a language's eventual disappearance.
That the 1931 census could, for the first time, capture this intermediate stage statistically was itself an intellectual advance in Indian linguistic study. The underlying trend it captured was one of slow attrition among small tribal languages under pressure from Hinduisation, migration and Hutton noted as “the ubiquitous increase of easy communications”.




Table: Mother-tongue speakers of selected major languages, 1931 (undivided India)
Language | Total mother-tongue speakers | Recorded as subsidiary language elsewhere |
Bengali | 53,037,841 | 339,048 |
Marathi | 20,789,201 | 10,39,414 |
Tamil | 20,177,594 | 755,118 |
Punjabi | 15,725,193 | 210,446 |
Kanarese (Kannada) | 11,196,103 | 1,378,517 |
Gujarati | 10,633,861 | 256,415 |
Oriya | 10,928,759 | 351,163 |
Malayalam | 9,101,440 | 183,829 |
Source: Census of India 1931, Table XV, Part II (“Bilingualism”). Figures for the composite Hindi/Urdu/Hindustani/Western Hindi returns are omitted here because, as discussed above, they were tabulated under contested and shifting categories that make a single clean figure misleading.
Case Study: New Alphabet in Chin Hills
The appendix of the 1931 census report documents a unique correspondence between Hutton and the Census Superintendent of Burma concerning a newly developed writing system in the Chin Hills. This script was created by Pau Chin Hau, a religious reformer of the Kamhow-Sokte people, to accompany a new religion.

Drawing visual inspiration from the Burmese and Roman alphabets, the script's early documentation in the report included a printed translation of the Sermon on the Mount and pages from a spelling book that illustrated consonant-vowel syllable signs and three distinct tonal marks.

Although Hutton was initially skeptical of the script's utility, fearing its vast number of monosyllabic characters made it too complex to record. The local administrators reported a more successful adaptation. Pau Chin Hau and his followers quickly consolidated the system, reducing it to 21 consonants alongside standard vowels and several Burmese-like tones.
The project was carried forward by Pau Chin Hau's son, Pow Ko Chin, and a schoolteacher named Than Chin Kham, who began translating the Gospel of Matthew into the new characters using an existing Roman-script Chin New Testament.
Ultimately, this case study highlights a broader trend within the 1931 census. It demonstrates that language in the region was not merely a static variable, but was actively being shaped, contested, and in some cases, literally created from scratch.
Linguistic Identity Formation
It is worth stepping back from the individual controversies to note what scholars such as Sarangi, building on Bernard Cohn's foundational work on caste and religion in colonial census-making, argue was happening at a structural level. Enumeration, according to Sarangi, was never simply a passive record of a pre-existing linguistic reality. It was itself a technology of classification that helped to produce the very communities it purported to count (Sarangi, 2009).
Census categories such as “mother-tongue,” “commonly spoken language,” or “Hindustani” were, in Sarangi's phrase, part of a “logic of numbers” that further intensified and politicized the linguistic-political struggle.
Ian Hacking's observation that “counting is hungry for categories,” and that many of the social categories used to describe people are themselves by-products of the needs of enumeration, applies with particular force to colonial India's language returns (Sarangi, 2009).
This dynamic was reinforced by the way census language data fed directly into constitutional and administrative processes with real distributive stakes. The 1931 census coincided with the run-up to the Government of India Act of 1935 and the associated debates over communal and provincial representation.
Hutton noted that even before the census tabulation was fully completed, both the Franchise Committee and the Orissa Boundary Committee had already requested the preliminary figures to help determine voting rights and outline electoral districts.
In this sense the census was not merely reporting on linguistic India but it was actively supplying the evidentiary basis on which provincial boundaries, educational funding, and eventually electoral representation would be redrawn.
This was precisely why communities on every side of every linguistic frontier had such strong incentives to inflate, understate or otherwise shape their own returns (Hutton, 1933, Sarangi, 2009).
Contemporary and Later Implications
In the immediate term, the 1931 language data helped settle one major boundary question. It was central to the evidence base for the creation of the separate province of Orissa on 1 April 1936, the first linguistically defined province carved out under British rule. For this a comparative table tracking the growth of Oriya speakers on Ganjam plains between 1881 to 1931 was prepared. It was hoped to serve as the objective correction to the exaggerated figures claimed by competing factions.
During this fifty-year period, the total population of the plains grew from 1,500,301 to 2,053,381, while the number of primary Oriya speakers rose from 748,904 (50 percent) to 1,079,337 (53 percent). When incorporating those who spoke Oriya as a secondary language, the total number of people with a command of the language reached 1,184,909, representing 57.7 percent of the local population.
The 1930s, as Sarangi notes, were more broadly “an era of linguistic nationalism in colonial India,” with parallel demands surfacing around Telugu–Tamil boundaries in Madras, Marathi–Gujarati divisions in Bombay Presidency, and the broader Bihari-versus-Bengali question that had already produced the separate province of Bihar and Orissa in 1912.
The Hindi–Hindustani–Urdu question the 1931 census tried and failed to defuse did not go away either. It resurfaced with full force in the Constituent Assembly's language debates of 1949. It eventually led to a constitutional compromise that made Hindi in the Devanagari script the official language of the Union, while keeping English as an associate official language. Over time, twenty-one other languages were added to the Eighth Schedule alongside Hindi, and Hindi-Urdu identity remains politically significant in both independent India and Pakistan.
The transformation of Urdu into a marker of Muslim identity and Pakistani nationhood, and of Sanskritised Hindi into a marker of Hindu-majoritarian nationhood, has clear roots traceable through the census categories this essay has described.
The tribal-language material in the 1931 chapter has its own afterlife. The pattern of gradual bilingual transition and eventual language shift that Hutton's contributors documented so carefully anticipated, by several decades, what linguists today call language endangerment — a framework now applied to many of the same Munda, Dravidian-tribal and Tibeto-Burman languages named in the 1931 report, several of which (Birhor and Korwa among them) remain listed as endangered, some critically so, in UNESCO's Atlas of the World's Languages in Danger. The 1931 census is, in that sense, an early and unusually granular data point in a story of linguistic loss that continues in India today.
Finally, similar to caste returns, the very unavailability of comparable later data gives the 1931 returns an outsized historiographical importance. Later Indian censuses since 1961 did not record bilingualism as thoroughly, and the Partition of India changed the population, making direct comparisons difficult. As a result, the 1931 records offer the most complete picture of language use, bilingualism, and tribal languages across undivided India just before the final phase of the independence movement.
(Authors: Gaurav Kalyani works as Research Associate for the India State Stories Project at the Center for Legislative Education and Research, FLAME University, Pune;
Shivakumar Jolad works as Associate Professor (Public Policy), and is the Chair of Center for Legislative Education and Research and Director India State Stories, FLAME University, Pune)
(Author Contributions: Gaurav did primary research and writing. Shivakumar contributed to conceptualization, research and editing)
Bibliography
Hutton, J. H. (1933). Census of India, 1931: Volume I — India, Part I: Report, Chapter X - Languages. Office of the Census Commissioner.
Hutton, J. H. (1933). Census of India, 1931: Volume I — India, Part II: Tables. Office of the Census Commissioner.
Hutton, J. H. (1932). The Indian census of 1931. Journal of the Royal Society of Arts, 80(4157), 785–797.
Shirras, G. F. (1935). The census of India, 1931. The Geographical Review, 25(3), 434–450.
Sarangi, A. (2009). Enumeration and the linguistic identity formation in colonial North India. Studies in History, 25(2), 197–227.




Comments