AI Literacy:

How LLMs Work, Their Biases, and Their Operation

Authors

DOI:

https://doi.org/10.61454/yhetdk61

Keywords:

AI literacy, large language models, tokenization, algorithmic bias, RLHF, benchmarks, corpus transparency

Abstract

Large Language Models (LLMs) have moved from a specialized object of computational linguistics to a routine infrastructure for professional, scientific and administrative production. This essay proposes an AI literacy understood as the technical-epistemic comprehension of the LLM artifact, organized in three axes: how it works, what biases it inscribes, and what responsible operability can be expected from its use. From a qualitative-interpretive paradigm, a documentary research design with a critical-analytical approach is adopted, following Bowen’s (2009) model, drawing on primary literature indexed in Web of Science, Scopus, Google Scholar, arXiv and the ACL Anthology between 2017 and May 2026, on transformer architectures (Vaswani et al., 2017), tokenization (Petrov et al., 2023; Ahia et al., 2023), reinforcement learning from human feedback (Ouyang et al., 2022; Casper et al., 2023), construct validity (Raji et al., 2021; Bowman and Dahl, 2021), and data coloniality (Quijano, 2000; Couldry and Mejias, 2019). The analysis identifies three central technical operations (tokenization, attention and next-token prediction) whose design decisions penalize Spanish relative to English with a tokenization premium of 1.55 to 1 in GPT-3.5 and GPT-4; three structural layers of bias inscription (corpus, RLHF alignment and benchmarks) that invalidate the neutrality claim of commercial systems; a drop in aggregate transparency of foundation models from 58 to 40.69 points out of 100 between 2024 and 2025; and a gap of up to 30.2 percentage points between the best and worst languages in MMLU-ProX. The article concludes that technical-epistemic literacy is a necessary condition for Latin American academia, public bodies, private sector, and the third sector to articulate situated, transparent, and sovereign AI policies.

Downloads

Download data is not yet available.

References

Ahia, O., Kumar, S., Gonen, H., Kasai, J., Mortensen, D. R., Smith, N. A., y Tsvetkov, Y. (2023). Do all languages cost the same? Tokenization in the era of commercial language models. En Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 9904-9923). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2023.emnlp-main.614

Atari, M., Xue, M. J., Park, P. S., Blasi, D. E., y Henrich, J. (2023). Which humans? PsyArXiv. https://doi.org/10.31234/osf.io/5b26t DOI: https://doi.org/10.31234/osf.io/5b26t

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv. https://doi.org/10.48550/arXiv.2212.08073

Balloccu, S., Schmidtová, P., Lango, M., y Dušek, O. (2024). Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs. En Proceedings of the 18th Conference of the European Chapter of the ACL (Vol. 1, pp. 67-93). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.eacl-long.5

Bender, E. M., Gebru, T., McMillan-Major, A., y Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? En Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922 DOI: https://doi.org/10.1145/3442188.3445922

Bender, E. M., y Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. En Proceedings of the 58th Annual Meeting of the ACL (pp. 5185-5198). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.463 DOI: https://doi.org/10.18653/v1/2020.acl-main.463

Birhane, A. (2020). Algorithmic colonization of Africa. SCRIPTed, 17(2), 389-409. https://doi.org/10.2966/scrip.170220.389 DOI: https://doi.org/10.2966/scrip.170220.389

Bommasani, R., Klyman, K., Longpre, S., Kapoor, S., Liang, P., et al. (2025). Foundation Model Transparency Index 2025. arXiv. https://doi.org/10.48550/arXiv.2512.10169

Bowen, G. A. (2009). Document analysis as a qualitative research method. Qualitative Research Journal, 9(2), 27-40. https://doi.org/10.3316/QRJ0902027 DOI: https://doi.org/10.3316/QRJ0902027

Bowman, S. R., y Dahl, G. E. (2021). What will it take to fix benchmarking in natural language understanding? En Proceedings of the 2021 Conference of the North American Chapter of the ACL: Human Language Technologies (pp. 4843-4855). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.385 DOI: https://doi.org/10.18653/v1/2021.naacl-main.385

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020). Language models are few-shot learners. En Advances in Neural Information Processing Systems, 33 (pp. 1877-1901). Curran Associates.

Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., et al. (2023). Open problems and fundamental limitations of Reinforcement Learning from Human Feedback. arXiv. https://doi.org/10.48550/arXiv.2307.15217

Couldry, N., y Mejias, U. A. (2019). Data colonialism: Rethinking big data’s relation to the contemporary subject. Television and New Media, 20(4), 336-349. https://doi.org/10.1177/1527476418796632 DOI: https://doi.org/10.1177/1527476418796632

Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press. DOI: https://doi.org/10.12987/9780300252392

Cronbach, L. J., y Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957 DOI: https://doi.org/10.1037/h0040957

Denzin, N. K., y Lincoln, Y. S. (Coords.). (2012). Manual de investigación cualitativa: Vol. IV. Métodos de recolección y análisis de datos. Editorial Gedisa.

Devlin, J., Chang, M.-W., Lee, K., y Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. En Proceedings of the 2019 Conference of the NAACL: HLT (Vol. 1, pp. 4171-4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423 DOI: https://doi.org/10.18653/v1/N19-1423

Durmus, E., Nyugen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., et al. (2023). Towards measuring the representation of subjective global opinions in language models. arXiv. https://doi.org/10.48550/arXiv.2306.16388

Floridi, L. (2013). The ethics of information. Oxford University Press. DOI: https://doi.org/10.1093/acprof:oso/9780199641321.001.0001

Floridi, L. (2023). AI as agency without intelligence: On ChatGPT, large language models, and other generative models. Philosophy and Technology, 36(15). https://doi.org/10.1007/s13347-023-00621-y DOI: https://doi.org/10.1007/s13347-023-00621-y

Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., et al. (2024). Are we done with MMLU? arXiv. https://doi.org/10.48550/arXiv.2406.04127

Goffman, E. (1981). Forms of talk. University of Pennsylvania Press.

Goodfellow, I., Bengio, Y., y Courville, A. (2016). Deep learning. MIT Press.

Haraway, D. (1988). Situated knowledges: The science question in feminism and the privilege of partial perspective. Feminist Studies, 14(3), 575-599. https://doi.org/10.2307/3178066 DOI: https://doi.org/10.2307/3178066

Havaldar, S., Pressimone, S., Wong, E., y Ungar, L. H. (2023). Multilingual language models are not multicultural: A case study in emotion. En Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis (pp. 202-214). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2023.wassa-1.19

Henrich, J., Heine, S. J., y Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X DOI: https://doi.org/10.1017/S0140525X0999152X

Jackson, R., Liu, Y., y Patel, A. (2025). Omniscience Index: Measuring calibrated knowledge in large language models. Artificial Analysis Reports.

Joshi, P., Santy, S., Budhiraja, A., Bali, K., y Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. En Proceedings of the 58th Annual Meeting of the ACL (pp. 6282-6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560 DOI: https://doi.org/10.18653/v1/2020.acl-main.560

Krippendorff, K. (2018). Content analysis: An introduction to its methodology (4th ed.). SAGE Publications. DOI: https://doi.org/10.4135/9781071878781

Kudo, T., y Richardson, J. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. En Proceedings of the 2018 Conference on EMNLP: System Demonstrations (pp. 66-71). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-2012 DOI: https://doi.org/10.18653/v1/D18-2012

La Fontaine, G. (2024). Sobre loros estocásticos: Una mirada a los modelos grandes de lenguaje. Lógoi: Revista de Filosofía, 45, 75-87. DOI: https://doi.org/10.62876/lr.vi45.6480

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741 DOI: https://doi.org/10.1037/0003-066X.50.9.741

Mignolo, W. D. (2002). The geopolitics of knowledge and the colonial difference. South Atlantic Quarterly, 101(1), 57-96. https://doi.org/10.1215/00382876-101-1-57 DOI: https://doi.org/10.1215/00382876-101-1-57

Mitchell, M., y Krakauer, D. C. (2023). The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13), e2215907120. https://doi.org/10.1073/pnas.2215907120 DOI: https://doi.org/10.1073/pnas.2215907120

OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. (2022). Training language models to follow instructions with human feedback. En Advances in Neural Information Processing Systems, 35 (pp. 27730-27744). Curran Associates. DOI: https://doi.org/10.52202/068431-2011

Perrigo, B. (2023, 18 de enero). Exclusive: OpenAI used Kenyan workers on less than $2 per hour to make ChatGPT less toxic. TIME. https://time.com/6247678/openai-chatgpt-kenya-workers/

Petrov, A., La Malfa, E., Torr, P. H. S., y Bibi, A. (2023). Language model tokenizers introduce unfairness between languages. En Advances in Neural Information Processing Systems, 36 (pp. 36963-36990). Curran Associates. DOI: https://doi.org/10.52202/075280-1608

Pezeshkpour, P., y Hruschka, E. (2024). Large language models sensitivity to the order of options in multiple-choice questions. En Findings of the ACL: NAACL 2024 (pp. 2006-2017). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.findings-naacl.130

Quijano, A. (2000). Colonialidad del poder, eurocentrismo y América Latina. En E. Lander (Comp.), La colonialidad del saber: eurocentrismo y ciencias sociales. Perspectivas latinoamericanas (pp. 201-246). CLACSO.

Raji, I. D., Bender, E. M., Paullada, A., Denton, E., y Hanna, A. (2021). AI and the everything in the whole wide world benchmark. En Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (Vol. 1).

Ricaurte, P. (2022). Ethics for the majority world: AI and the question of violence at scale. Media, Culture and Society, 44(4), 726-745. https://doi.org/10.1177/01634437221099612 DOI: https://doi.org/10.1177/01634437221099612

Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386-408. https://doi.org/10.1037/h0042519 DOI: https://doi.org/10.1037/h0042519

Rumelhart, D. E., Hinton, G. E., y Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533-536. https://doi.org/10.1038/323533a0 DOI: https://doi.org/10.1038/323533a0

Sanderson, G. (2024). Attention in transformers, visually explained. 3Blue1Brown. https://www.3blue1brown.com/lessons/attention

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., y Hashimoto, T. (2023). Whose opinions do language models reflect? En Proceedings of the 40th International Conference on Machine Learning (pp. 29971-30004). PMLR.

Sennrich, R., Haddow, B., y Birch, A. (2016). Neural machine translation of rare words with subword units. En Proceedings of the 54th Annual Meeting of the ACL (Vol. 1, pp. 1715-1725). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162 DOI: https://doi.org/10.18653/v1/P16-1162

Singh, S., Romanou, A., Fourrier, C., Adelani, D. I., Ngoc Khai, J., Vila-Suero, D., et al. (2025). Global-MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation. En Proceedings of the 63rd Annual Meeting of the ACL (Vol. 1). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2025.acl-long.919

Teixeira, F., Almeida, R., y Silva, M. (2025). Tokenization, context windows, and the architecture of language models for educators. Computers and Education: Artificial Intelligence, 8, 100245.

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv. https://doi.org/10.48550/arXiv.2307.09288

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., y Polosukhin, I. (2017). Attention is all you need. En Advances in Neural Information Processing Systems, 30 (pp. 5998-6008). Curran Associates.

Veronelli, G. (2015). Sobre la colonialidad del lenguaje. Universitas Humanística, 81, 33-58. https://doi.org/10.11144/Javeriana.uh81.scl DOI: https://doi.org/10.11144/Javeriana.uh81.scdl

Xuan, W., Yang, K., Tan, Y., Tian, X., Qin, Y., Zhang, M., et al. (2025). MMLU-ProX: A multilingual benchmark for advanced large language model evaluation. arXiv. https://doi.org/10.48550/arXiv.2503.10497 DOI: https://doi.org/10.18653/v1/2025.emnlp-main.79

Published

2026-07-27

Issue

Section

Artículos

How to Cite

AI Literacy: : How LLMs Work, Their Biases, and Their Operation. (2026). Espectro Investigativo Latinoamericano, 8(2), 118-133. https://doi.org/10.61454/yhetdk61

Similar Articles

1-10 of 50

You may also start an advanced similarity search for this article.