Code-Switching in Social Media: A Corpus-Based Analysis of Urdu-English Mixing on Pakistani Digital Platforms
DOI:
https://doi.org/10.63954/WAJSS.4.2.55.2025Keywords:
social media corpora, corpus linguistics, Pakistani EnglishAbstract
New communicative spaces have emerged with the proliferation of social media platforms, where multilingual speakers mix languages in context-dependent, creative, and fluid ways. The present study investigates code switching (CS) between Urdu and English in social media (SM) corpora of Pakistan, based on a corpus of 12,000 posts that were gathered from different social media platforms (Facebook, Twitter/X, Instagram, and WhatsApp) from 2022 to 2024. The study uses corpus linguistics methodology (frequency analysis, concordance searches and KWIC analysis) to investigate the patterns of CS in informal digital registers. The study finds that all three phenomena – inter-sentential, intra-sentential and tag-switching – are attested, the latter being the most common. The most common categories that switch are nouns, verbs, and discourse markers. Key sociolinguistic motivations are foregrounded, including identity negotiation, humor, stance-marking, and code prestige. It also addresses methodological problems in creating Romanized Urdu corpora and proposes an approach to annotating multilingual social media data. The findings enrich the existing body of literature on Computer-Mediated Communication (CMC), digital multilingualism, and corpus-assisted sociolinguistics in the context of South Asia.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Arjumand Shaheen, Muhammad Kamran Abbas Ismail, Aimen Javed, Rimsha Bibi

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright and Licensing
Publication is open access
Creative Commons Attribution License - CC BY- 4.0
Copyrights: The author retains unrestricted copyrights and publishing rights
