25 | Vernacular Self-Description and Census Categories: Language in the 1911 Census
- Mar 25
- 13 min read
By Nishitha Mandava
Published on: 27 July 2026
Introduction
During the 1911 census, the enumerators were instructed to record ‘the language which each person ordinarily uses in his own home’(Gait, 1913). While this appeared to leave the answer to the person being enumerated, by the early 20th century, the census authorities already had their own understanding of what India’s languages were, how they were related to one another, and where their boundaries lay. When people’s answers did not correspond to these classifications, the census did not always accept their self-description as the correct one.
Across much of northern India, for instance, millions simply called their language Hindi. Of those who returned an Indo-Aryan language as their mother tongue, 82 million (more than a third), described it by that single name. But for census authorities this was not precise enough. Census classifications distinguished between Western Hindi, Eastern Hindi and Bihari, and further divided Bihari into Magahi, Maithili and Bhojpuri.
In the Central Provinces, speakers of Nimari and Malwi commonly described their language as Hindi, although the census classified both as dialects of Rajasthani. In the North-West Frontier Province, many people were returned as speaking Punjabi even though the report maintained that most actually spoke Lahnda, which it classified as a separate language belonging to a different linguistic group altogether (Gait, 1913).
These discrepancies reflected two different ways of naming language. By 1911, decades of census-taking and philological investigation had produced an elaborate official classification of India’s languages. Hence, the census authorities interpreted that a person who identified their language as Hindi might actually be speaking Bihari, Magahi, Maithili, Bhojpuri, or Nimari (Gait, 1913). Vernacular self-descriptions were thus assessed against a pre-existing classificatory scheme that was considered to be more ‘accurate.’
The 1911 census made this distinction explicit by presenting two versions of India’s linguistic population.
The first reproduced the languages actually entered in the census schedules. The second broke broad terms such as Hindi into what the report called their ‘proper constituents,’ based on the conclusions of the Linguistic Survey about the geographical distribution of different languages.
It also structured the relationship between provincial and imperial knowledge. For instance, the extract from the report below shows that the language Ahirwati was enumerated as a dialect of Western Hindi by officials at the provincial level but it was classified in the all-India census as Rajasthani based on Grierson’s linguistic scheme.
"4. Dialects have been dealt with in the manner indicated in the Index of Indian languages printed at pages 261—292 of the India Administrative Volume, 1901, with such modifications (suggested by Sir George Grierson) as have been found necessary in view of the progress since made by the Linguistic Survey. In a few cases, as noted below, this later information was not available in time for use in the Provincial Language Tables, and the classification here made differs from that adopted in the Provincial volumes”:-
Language | Classification in the Provincial Tables | Classification in the India Table |
Ahirwātī | Western Hindi | Rājasthāni. |
Bahrūpia | Gipsy | Ditto. |
Banjārī or Labhānī | Ditto | Ditto. |
Baorī | Ditto | Bhīl. |
Bhilālī | Rājasthāni | Ditto. |
Charanī | Gipsy | Ditto. |
Gujarī | Western Pahārī | Ditto. |
Labānī, Labānkī or Lamānī | Gipsy | Rājasthāni. |
Magh | Burmese | Arakanese. |
Malvī | Gujarātī | Rājasthāni. |
Panchālī | Marāthī | Bhīl. |
Pardhī or Takankarī | Gipsy | Ditto. |
Siyālgirī | Gujarātī | Ditto. |
Western Pahārī | Western Group | Northern Group. |
Note: The term "Ditto" refers to the classification immediately above it in the same column.
(Source: Gait, 1913)
These revisions reveal that linguistic classification was no longer determined by provincial administrative practice but increasingly by the philological authority of the Linguistic Survey.
Unlike the 1901 census, which had reassigned persons reporting their language as ‘Hindi’ to Western Hindi, Eastern Hindi, or Bihari based on the birthplace of the individual, the 1911 report retained such responses under the broad category of Hindi in its general tables, while recording more detailed linguistic classifications in its subsidiary analyses. The 1911 census sought to standardise the multitude of vernacular names through which languages were reported. It provided tables showing how locally used designations were to be entered under official census names, for instance, Andhra as Telugu, Talaing as Mon, and Oraon as Kurukh (Gait, 1913).
After decades of enumeration and ethnographic and philological research, colonial officials had developed their systems of classification for categorising India’s social groups and their relations to one another. But the identities and affiliations reported by people did not always neatly fit into these categories through which census officials had come to understand Indian society. The tension arose when vernacular self-description did not conform to that official knowledge.
While the censuses of the preceding decades repeatedly emphasised the difficulty of classifying India’s social identities, pointing to their complexity and local variation, we observe a shift in the 1911 census. By this point, colonial officials had developed an elaborate classificatory framework through decades of enumeration and surveying.
As a result, vernacular self-description was no longer treated as authoritative evidence of social identity but as material requiring interpretation based on the accumulated body of colonial knowledge. This essay in the subsequent sections studies the tensions produced by this shift wherein competing understandings of language co-existed and often clashed.
The Geography of Languages
The Indo-European (Aryan) languages, with 232.8 million speakers or 74.3 percent of the population was dominant everywhere except Burma, the Assam hills, and the peninsula south of a line running roughly from Kolhapur to Puri. South of that line the Dravidian family which comprised of nearly 63 millions or one-fifth of the total and in some outlier regions including Central Provinces (territories of present-day Madhya Pradesh), the Chota Nagpur plateau (Kurukh, Malto), and, most mysteriously, Brahui in distant Baluchistan, ‘one of the greatest riddles in Indian philology’ (Gait, 1913).
The Tibeto-Chinese family, though spread over the whole Himalayan arc from Ladakh to the Mishmi country and over Burma included 13 millions, about 4 percent. The Austro-Asiatic family, with 4.4 millions largely in Chota Nagpur. Its Munda dialects were probably spoken across much of the Indo-Gangetic plain before the arrival of the Aryans.
Although they later disappeared from most of the region, traces of them remain in the pronominalised dialects of the Himalayas and possibly in the verb forms of the Bihari languages. Similarly, Mon-Khmer languages were once spoken across much of mainland Southeast Asia before the spread of Tibeto-Burman and Indo-Aryan languages.

The Dravidian family fell into three divisions. The Andhra group (24.1 million) was predominantly Telugu (23.5 million), spoken north of Madras city and in eastern Hyderabad, the second language of India by count.
The Dravida group (37.1 million) comprised Tamil (18.1 million) in the centre and south-east of Madras, Kanarese (10.5 million) in Mysore, southern Hyderabad and the Canara districts, Malayalam (6.8 million) on the west coast, Tulu (0.6 million) in South Canara, and the northern outliers Kurukh (0.8 million) and Brahui. Between them lay the intermediate Gondi group (1.5 million) of the central hills (Gait, 1913).
Unlike the tribal tongues of the north, the great Dravidian languages were yielding nothing to Aryan speech: “Kanarese is not being pushed back by Marathi, nor are Telugu and Tamil yielding to Aryan tongues” (Gait, 1913). Within the Tibeto-Chinese language family, Burmese was by far the largest language, with around 8 million speakers. It was much larger than the many smaller languages in the same family, including Arakanese, Manipuri, Bodo, Garo, Kachin, and the Chin and Naga dialects.
Another branch of the broader family, the Siamese-Chinese sub-family, included Karen (1.1 million speakers) and Shan (0.9 million speakers). In the Austro-Asiatic language family, the largest language group in India was Kherwari, with around 3.4 million speakers. Its most widely spoken dialect, Santali, had about 2.1 million speakers and remained one of India's strongest and most widely used tribal languages (Gait, 1913).
Re-Classifying Languages
Edward Albert Gait, the 1911 Census Commissioner, in his report writes that the mother-tongue of 313.5 million people was recorded. Mother tongue here was defined as ‘the language which each person ordinarily uses in his own home’(Gait, 1913).
The languages were classified under the scheme drawn up by Sir George Grierson, Director of the Linguistic Survey of India, whose linguistic survey was then transforming how languages were being classified in the Indian subcontinent. The following table shows how languages were classified in the census based on Grierson’s classifications which divided languages into various linguistic families:
Family / Sub-family / Branch | Principal languages and groups | Speakers |
A. VERNACULARS OF INDIA (c. 220 languages) | 312,948,881 | |
Indo-European Family (Aryan Sub-family) | 232,822,511 | |
Eranian Branch (Eastern group) | Pashto (1,554,465); Baloch (504,586) | 2,066,654 |
Indian Branch — Pisacha (non-Sanskritic) sub-branch | Kashmiri (1,180,632); Shina, Khowar, Kohistani | 1,207,159 |
Indian Branch — Sanskritic sub-branch | 229,548,098 | |
North-Western group | Lahnda (4,779,138 + Siraiki 206,452); Sindhi (3,669,935 + Kachchhi 389,736) | 8,449,073 |
Southern group | Marathi (19,806,636); Konkani (495,902); Singhalese | 20,330,838 |
Eastern group | Bengali (48,367,915); Oriya (10,162,321); Assamese (1,533,822); Bihari (as returned, 398,294) | 60,462,352 |
Mediate group | 'Hindi' as returned (82,003,235); Eastern Hindi (2,423,392) | 84,426,627 |
Western group | Panjabi (15,876,758); Rajasthani (14,067,590); Western Hindi (14,037,882); Gujarati (10,682,248); Bhil languages (1,435,445) | 54,664,178 |
Northern group | Western Pahari (1,526,475); Eastern Pahari/Naipali (208,932); Central Pahari | 1,739,172 |
Dravidian Family | 62,718,961 | |
Dravida group | Tamil (18,128,365); Kanarese (10,525,739); Malayalam (6,792,277); Kurukh/Oraon (800,328); Tulu (563,453); Brahui (174,229) | 37,094,393 |
Intermediate languages | Gondi, etc. | 1,527,157 |
Andhra group | Telugu (23,542,861); Kandh/Kui (530,476); Kolami | 24,097,411 |
Tibeto-Chinese Family | 12,972,512 | |
Tibeto-Burman sub-family | Burmese (c. 8,000,000); Arakanese, Manipuri, Bodo (c. 300,000 each); Tibetan, Garo, Kachin, Chin dialects | 10,932,775 |
Siamese-Chinese sub-family | Karen (c. 1,100,000); Shan (c. 900,000) | 2,039,737 |
Austro-Asiatic Family | 4,398,640 | |
Mon-Khmer sub-family | Khasi (200,872); Mon/Talaing (179,444); Palaung-Wa (166,683); Nicobarese (8,418) | 555,417 |
Munda sub-family | Kherwari (3,357,661: Santali 2,138,015, Mundari 599,580, Ho 420,108); Savara (166,280); Kurku (136,909); Kharia (126,583) | 3,843,223 |
Malayo-Polynesian Family | Selung, Malay | 6,179 |
Unclassified | Gipsy languages (28,294); Andamanese (1,324) | 29,618 |
B. VERNACULARS OF OTHER ASIATIC COUNTRIES & AFRICA | Chinese (113,450); Persian (56,589); Arabic (42,102) | 223,110 |
C. EUROPEAN LANGUAGES | English chief among them | 321,224 |
INDIA | 313,493,215 |
The Grierson classification of Indian languages as adopted in the census of 1911, with speakers of each division (Source: Gait, 1913). Figures 'as returned' reflect the census schedules, not the Linguistic Survey's corrections.

The 1911 census followed Grierson’s revised scheme of the Linguistic Survey, and this produced an important change: the re-affiliation of the Munda languages. In the 1901 census, for which Grierson authored the chapter on language, Munda and Dravidian had been treated as sub-families of the larger ‘Dravido-Munda’ family (Majeed, 2019). However, later developments in his research made him conclude that the two groups were not closely related.
By comparing their sound systems, grammar, noun classification, methods of counting, and verb structures, he concluded that they had fundamentally different linguistic characteristics and therefore ‘the two groups of languages have no real connection’ (Gait, 1913).

Based on these findings, the Munda languages were reclassified as part of a new Austro-Asiatic language family. Building on the work of the linguist Wilhelm Schmidt, Grierson suggested that Munda languages such as Santali and Mundari were more closely related to languages spoken across eastern India, Burma, and Southeast Asia, including Mon, Khasi, Nicobarese, and Palaung-Wa.


Schmidt further suggested that Austro-Asiatic itself formed part of an even larger ‘Austric’ family extending across much of Asia and the Pacific, although this broader theory remains contested today.
This reclassification also led to other changes in the census. Kashmiri was transferred to a different branch of the Indo-Aryan languages and Brahui was confirmed as a Dravidian language despite its geographical isolation from other Dravidian languages.

Several Munda dialects were consolidated under a single language called Kherwari, and newly documented languages from Burma and northeastern India were provisionally incorporated into the revised linguistic scheme. These changes reflect the growing reliance of the census on linguistic experts rather than simply relying on how people themselves identified their languages.
These changing language classifications also had implications for Herbert Risley’s race theory that was dependent on anthropometric measurements which we discussed in the context of the 1901 census.
While Grierson had demonstrated that Munda and Dravidian languages belonged to entirely different linguistic families, Risley's anthropometric measurements had found little physical difference between Munda and Dravidian speaking populations. Grierson speculated that the physical type conventionally described as ‘Dravidian’ might in fact be associated with Munda-speaking peoples.
Interestingly, Gait rejected this conclusion, arguing instead that the Dravidian peoples most likely originated in southern India and that the presence of the Brahui language in Baluchistan represented an isolated case of linguistic diffusion rather than evidence of a common racial origin (Gait, 1913). Through this, he cautioned against using linguistic evidence for drawing inferences on racial descent.
Political Controversies and ‘Decaying’ Languages
The census found itself embroiled in various political controversies. One of them was the Hindi-Urdu political controversy which was at its height in the early 20th century. The report mentions that among some educated Hindus, there was a tendency to minimise linguistic differences within northern India and through this they would claim that there was only one widely understood Hindi language.
At the same time, some Muslims maintained that Urdu was the language not only of their co-religionists but also of many Hindus in northern India. These competing claims sometimes affected the entries made in census schedules, particularly in the Punjab and the United Provinces (present-day Uttar Pradesh).
The problem was especially prevalent in the United Provinces. The report traced the controversy back to government orders issued in 1900 permitting court documents to be written in either script and, in some cases, requiring both. Although the original question concerned script rather than spoken language, it became connected to the Hindi-Urdu controversy.
A further dispute arose when primary school textbooks were revised in 1910. An attempt was made to differentiate textbooks written in the two scripts by making one a vehicle for Persianised Urdu and the other for Sanskritised or High Hindi. The census report complained that this dispute over scripts, literary styles and school textbooks was again transferred to the question of spoken language (Gait, 1913).
The census authorities received complaints that Hindu enumerators were recording Hindi regardless of the answers given by those they enumerated, while Muslim enumerators were accused of doing the same for Urdu. The Superintendent of Census in the United Provinces believed that such cases had occurred, particularly in cities where the controversy was most intense.
The resulting figures showed Urdu declining by one-fifth since 1901, while individual districts displayed what the report described as ‘absurd differences.’ The Superintendent concluded that the district figures were evidence only of the relative ‘strength or weakness of the agitation’ in particular places and that the controversy had ‘utterly falsified’ a set of statistics that was intended to measure something different (Gait, 1913).
Similar problems appeared elsewhere. In the Goalpara district of Assam, many people initially returned as Assamese speakers were subsequently classified as Bengali speakers by Bengali Charge Superintendents. A later local inquiry concluded that Assamese speakers were at least 30,000, or 35 percent, more numerous than the published census figures indicated.
The report also described the difficulty of classifying former tea-garden labourers who had settled permanently in Assam. Many had ancestral Munda or Dravidian languages but had developed a mixed form of speech containing Hindi, Bengali and Assamese. Assamese enumerators often classified this speech as Bengali simply because they knew it was not Assamese. In the Madras Presidency, meanwhile, the report attributed an apparent fall of 316,000 Oriya speakers in Ganjam partly to earlier Telugu-Oriya disputes that had led to deliberate misrepresentation by some enumerators in 1901.
Hence, the 1911 census was openly doubtful about the value of some of its own linguistic statistics, particularly for the Indo-Aryan languages. It nevertheless considered the returns for tribal languages more useful, especially for tracing whether particular languages were maintaining themselves or being replaced by other forms of speech. This distinction led the census to classify languages as either ‘dominant’ or ‘decaying’ and to examine the extent to which tribal languages were being abandoned (Gait, 1913).

In the Central Provinces and Berar, for example, the census found that Hindi and Marathi had replaced Gondi among more than half of the Gond population. Of nearly two and a half million Gonds, fewer than one and a half million were recorded as speaking Gondi.
Korku had declined less sharply, although 18,000 of 152,000 Korkus had abandoned their mother tongue. Among the Korwas, fewer than half retained their own language (Gait, 1913).
The report also observed that languages could survive in place names even after they had largely disappeared from everyday speech, identifying numerous villages, hills and rivers whose names it believed had been derived from Gondi.
Racial Descent and Language
In line with Grierson’s insights in the 1901 census, Gait also warned against treating language as a reliable indication of racial origin. The 1911 census emphasised that communities could abandon one language and adopt another within a short span of time. The Turungs of eastern Assam, for example, had discarded their earlier Shan language and adopted Singpho after spending several years captive among the Singhpo and were subsequently beginning to adopt Assamese.
Some Oraons near Ranchi had forgotten their own language and adopted a form of Mundari, which they again began to replace with Sadani. In Manipur, some Nagas and Kukis who became Hindus had also learnt Manipuri, while the report observed that hill people who moved into the plains of Burma could become Shan or Burmese within a generation (Gait, 1913).
The example of the Upper Chindwin district is further instructive on how rapidly language and cultural identification could change. The inhabitants spoke Kachin, wore Kachin dress and were called Kachins, but had begun learning Shan. Their headman stated that the community had earlier migrated from Assam, where its members had worn different clothes and spoken a language whose name they had entirely forgotten. The district account concluded that within two generations, they had lost almost all traces of their earlier origin.
It similarly argues that calling the population of the Upper Chindwin Burmese or Shan often meant only that their ancestors had, at some point, adopted Burmese or Shan. As the report puts it, ‘The Burmese language is the result of the Burmese domination. The Shan language is the result of the Shan domination’ (Gait, 1913).
On this basis, the 1911 Census explicitly cautioned against constructing racial theories from linguistic evidence.
A language spoken by a small group and surrounded by a more dominant language might preserve evidence of an earlier linguistic distribution, but the report argued that it could not safely be treated as proof of the racial origins of its speakers.
The language itself might once have displaced still earlier forms of speech. As Gait observed, ‘The dying languages of to-day were the dominant languages of a previous epoch’ (Gait, 1913).
Conclusion
In 1911, the census was engaged in two complementary classificatory operations. Broad vernacular labels were analytically disaggregated into what comparative philology regarded as their constituent languages, while multiple local names for the same language were also recorded and consolidated under a single standardised designation. Instead of only recording linguistic self-description, the census chose to regulate the vocabulary through which language itself could be known and enumerated.
This marked a significant shift from the 1901 census. Along with decades of surveying and Grierson’s Linguistic Survey of India furnished the colonial administrators with taxonomy for classification of languages, allowing them to claim authority over colloquial usage. Because of this, self-descriptions by the people were not always accepted as authoritative. Instead, their responses were further rationalised and re-classified by the officials, demonstrating the growing body of colonial knowledge of the colony.
But this ambition of classifying languages with philological precision encountered many hurdles including the Hindi-Urdu controversy. The local enumerators were not immune to such political contestations over language and this affected their returns, which the colonial officials were again left to rationalise and re-interpret.
Equally, the census’s own discussions of multilingualism, migration, language shift, rapid adoption and abandonment of languages complicated any straightforward relationship between language, territory, or community. This is especially evident in the case of the reclassification of the Munda languages into the Austro-Asiatic family.
Overall, the 1911 census becomes significant for understanding the ambition and limits of colonial knowledge. It represents a moment when linguistic classification became centralised under the authority of philological expertise, but it also emphasises the persistent instability of the social realities it set out to record.
(Authors: Nishitha Mandava, is an independent researcher and research consultant for
Center for Legislative Education and Research, FLAME University, Pune)
Sources:
Gait, E. A. (1913). Census of India, 1911 (Vol. 1, Pt. 1, Report). Superintendent Government Printing, India.
Gait, E. A. (1913). Census of India, 1911 (Vol. 1, Pt. II, Tables). Superintendent Government Printing, India.
Majeed, J. (2019). Colonialism and knowledge in Grierson’s linguistic survey of India. Routledge.




Comments