Literacidad en IA

Funcionamiento, Sesgos y Operatividad de los LLMs

Autores/as

DOI:

https://doi.org/10.61454/yhetdk61

Palabras clave:

literacidad en IA, modelos grandes de lenguaje, tokenización, sesgo algorítmico, RLHF, benchmarks, transparencia de datos

Resumen

Los modelos grandes de lenguaje (LLMs) han pasado de objeto de la lingüística computacional a infraestructura cotidiana de la producción profesional, científica y administrativa. Este ensayo propone una literacidad en IA entendida como comprensión técnico-epistémica del artefacto LLM, organizada en tres ejes: cómo funciona, qué sesgos inscribe y qué operatividad responsable cabe esperar de su uso. Desde un paradigma cualitativo-interpretativo, se adopta un diseño de investigación documental con enfoque crítico-analítico, siguiendo el modelo de Bowen (2009), a partir de literatura primaria indexada en Web of Science, Scopus, Google Scholar, arXiv y la ACL Anthology entre 2017 y mayo de 2026, sobre arquitecturas transformer (Vaswani et al., 2017), tokenización (Petrov et al., 2023; Ahia et al., 2023), aprendizaje por refuerzo con retroalimentación humana (Ouyang et al., 2022; Casper et al., 2023), validez de constructo (Raji et al., 2021; Bowman y Dahl, 2021) y colonialidad de datos (Quijano, 2000; Couldry y Mejias, 2019). El análisis identifica tres operaciones técnicas centrales (tokenización, atención y predicción del siguiente token) cuyas decisiones de diseño penalizan al español frente al inglés con un premium de tokenización de 1,55 a 1 en GPT-3.5 y GPT-4; tres capas estructurales de inscripción de sesgos (corpus, alineamiento por RLHF y benchmarks) que invalidan la pretensión de neutralidad de los sistemas comerciales; una caída de la transparencia agregada de los modelos de fundación de 58 a 40,69 puntos sobre 100 entre 2024 y 2025; y una brecha de hasta 30,2 puntos porcentuales entre la mejor y la peor lengua en MMLU-ProX. Se concluye que la literacidad técnico-epistémica es condición necesaria para que la academia, los organismos públicos, el sector privado y el tercer sector latinoamericanos articulen políticas de IA situadas, transparentes y soberanas.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

Ahia, O., Kumar, S., Gonen, H., Kasai, J., Mortensen, D. R., Smith, N. A., y Tsvetkov, Y. (2023). Do all languages cost the same? Tokenization in the era of commercial language models. En Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 9904-9923). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2023.emnlp-main.614

Atari, M., Xue, M. J., Park, P. S., Blasi, D. E., y Henrich, J. (2023). Which humans? PsyArXiv. https://doi.org/10.31234/osf.io/5b26t DOI: https://doi.org/10.31234/osf.io/5b26t

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv. https://doi.org/10.48550/arXiv.2212.08073

Balloccu, S., Schmidtová, P., Lango, M., y Dušek, O. (2024). Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs. En Proceedings of the 18th Conference of the European Chapter of the ACL (Vol. 1, pp. 67-93). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.eacl-long.5

Bender, E. M., Gebru, T., McMillan-Major, A., y Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? En Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922 DOI: https://doi.org/10.1145/3442188.3445922

Bender, E. M., y Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. En Proceedings of the 58th Annual Meeting of the ACL (pp. 5185-5198). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.463 DOI: https://doi.org/10.18653/v1/2020.acl-main.463

Birhane, A. (2020). Algorithmic colonization of Africa. SCRIPTed, 17(2), 389-409. https://doi.org/10.2966/scrip.170220.389 DOI: https://doi.org/10.2966/scrip.170220.389

Bommasani, R., Klyman, K., Longpre, S., Kapoor, S., Liang, P., et al. (2025). Foundation Model Transparency Index 2025. arXiv. https://doi.org/10.48550/arXiv.2512.10169

Bowen, G. A. (2009). Document analysis as a qualitative research method. Qualitative Research Journal, 9(2), 27-40. https://doi.org/10.3316/QRJ0902027 DOI: https://doi.org/10.3316/QRJ0902027

Bowman, S. R., y Dahl, G. E. (2021). What will it take to fix benchmarking in natural language understanding? En Proceedings of the 2021 Conference of the North American Chapter of the ACL: Human Language Technologies (pp. 4843-4855). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.385 DOI: https://doi.org/10.18653/v1/2021.naacl-main.385

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020). Language models are few-shot learners. En Advances in Neural Information Processing Systems, 33 (pp. 1877-1901). Curran Associates.

Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., et al. (2023). Open problems and fundamental limitations of Reinforcement Learning from Human Feedback. arXiv. https://doi.org/10.48550/arXiv.2307.15217

Couldry, N., y Mejias, U. A. (2019). Data colonialism: Rethinking big data’s relation to the contemporary subject. Television and New Media, 20(4), 336-349. https://doi.org/10.1177/1527476418796632 DOI: https://doi.org/10.1177/1527476418796632

Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press. DOI: https://doi.org/10.12987/9780300252392

Cronbach, L. J., y Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957 DOI: https://doi.org/10.1037/h0040957

Denzin, N. K., y Lincoln, Y. S. (Coords.). (2012). Manual de investigación cualitativa: Vol. IV. Métodos de recolección y análisis de datos. Editorial Gedisa.

Devlin, J., Chang, M.-W., Lee, K., y Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. En Proceedings of the 2019 Conference of the NAACL: HLT (Vol. 1, pp. 4171-4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423 DOI: https://doi.org/10.18653/v1/N19-1423

Durmus, E., Nyugen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., et al. (2023). Towards measuring the representation of subjective global opinions in language models. arXiv. https://doi.org/10.48550/arXiv.2306.16388

Floridi, L. (2013). The ethics of information. Oxford University Press. DOI: https://doi.org/10.1093/acprof:oso/9780199641321.001.0001

Floridi, L. (2023). AI as agency without intelligence: On ChatGPT, large language models, and other generative models. Philosophy and Technology, 36(15). https://doi.org/10.1007/s13347-023-00621-y DOI: https://doi.org/10.1007/s13347-023-00621-y

Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., et al. (2024). Are we done with MMLU? arXiv. https://doi.org/10.48550/arXiv.2406.04127

Goffman, E. (1981). Forms of talk. University of Pennsylvania Press.

Goodfellow, I., Bengio, Y., y Courville, A. (2016). Deep learning. MIT Press.

Haraway, D. (1988). Situated knowledges: The science question in feminism and the privilege of partial perspective. Feminist Studies, 14(3), 575-599. https://doi.org/10.2307/3178066 DOI: https://doi.org/10.2307/3178066

Havaldar, S., Pressimone, S., Wong, E., y Ungar, L. H. (2023). Multilingual language models are not multicultural: A case study in emotion. En Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis (pp. 202-214). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2023.wassa-1.19

Henrich, J., Heine, S. J., y Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X DOI: https://doi.org/10.1017/S0140525X0999152X

Jackson, R., Liu, Y., y Patel, A. (2025). Omniscience Index: Measuring calibrated knowledge in large language models. Artificial Analysis Reports.

Joshi, P., Santy, S., Budhiraja, A., Bali, K., y Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. En Proceedings of the 58th Annual Meeting of the ACL (pp. 6282-6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560 DOI: https://doi.org/10.18653/v1/2020.acl-main.560

Krippendorff, K. (2018). Content analysis: An introduction to its methodology (4th ed.). SAGE Publications. DOI: https://doi.org/10.4135/9781071878781

Kudo, T., y Richardson, J. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. En Proceedings of the 2018 Conference on EMNLP: System Demonstrations (pp. 66-71). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-2012 DOI: https://doi.org/10.18653/v1/D18-2012

La Fontaine, G. (2024). Sobre loros estocásticos: Una mirada a los modelos grandes de lenguaje. Lógoi: Revista de Filosofía, 45, 75-87. DOI: https://doi.org/10.62876/lr.vi45.6480

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741 DOI: https://doi.org/10.1037/0003-066X.50.9.741

Mignolo, W. D. (2002). The geopolitics of knowledge and the colonial difference. South Atlantic Quarterly, 101(1), 57-96. https://doi.org/10.1215/00382876-101-1-57 DOI: https://doi.org/10.1215/00382876-101-1-57

Mitchell, M., y Krakauer, D. C. (2023). The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13), e2215907120. https://doi.org/10.1073/pnas.2215907120 DOI: https://doi.org/10.1073/pnas.2215907120

OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. (2022). Training language models to follow instructions with human feedback. En Advances in Neural Information Processing Systems, 35 (pp. 27730-27744). Curran Associates. DOI: https://doi.org/10.52202/068431-2011

Perrigo, B. (2023, 18 de enero). Exclusive: OpenAI used Kenyan workers on less than $2 per hour to make ChatGPT less toxic. TIME. https://time.com/6247678/openai-chatgpt-kenya-workers/

Petrov, A., La Malfa, E., Torr, P. H. S., y Bibi, A. (2023). Language model tokenizers introduce unfairness between languages. En Advances in Neural Information Processing Systems, 36 (pp. 36963-36990). Curran Associates. DOI: https://doi.org/10.52202/075280-1608

Pezeshkpour, P., y Hruschka, E. (2024). Large language models sensitivity to the order of options in multiple-choice questions. En Findings of the ACL: NAACL 2024 (pp. 2006-2017). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.findings-naacl.130

Quijano, A. (2000). Colonialidad del poder, eurocentrismo y América Latina. En E. Lander (Comp.), La colonialidad del saber: eurocentrismo y ciencias sociales. Perspectivas latinoamericanas (pp. 201-246). CLACSO.

Raji, I. D., Bender, E. M., Paullada, A., Denton, E., y Hanna, A. (2021). AI and the everything in the whole wide world benchmark. En Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (Vol. 1).

Ricaurte, P. (2022). Ethics for the majority world: AI and the question of violence at scale. Media, Culture and Society, 44(4), 726-745. https://doi.org/10.1177/01634437221099612 DOI: https://doi.org/10.1177/01634437221099612

Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386-408. https://doi.org/10.1037/h0042519 DOI: https://doi.org/10.1037/h0042519

Rumelhart, D. E., Hinton, G. E., y Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533-536. https://doi.org/10.1038/323533a0 DOI: https://doi.org/10.1038/323533a0

Sanderson, G. (2024). Attention in transformers, visually explained. 3Blue1Brown. https://www.3blue1brown.com/lessons/attention

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., y Hashimoto, T. (2023). Whose opinions do language models reflect? En Proceedings of the 40th International Conference on Machine Learning (pp. 29971-30004). PMLR.

Sennrich, R., Haddow, B., y Birch, A. (2016). Neural machine translation of rare words with subword units. En Proceedings of the 54th Annual Meeting of the ACL (Vol. 1, pp. 1715-1725). Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162 DOI: https://doi.org/10.18653/v1/P16-1162

Singh, S., Romanou, A., Fourrier, C., Adelani, D. I., Ngoc Khai, J., Vila-Suero, D., et al. (2025). Global-MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation. En Proceedings of the 63rd Annual Meeting of the ACL (Vol. 1). Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2025.acl-long.919

Teixeira, F., Almeida, R., y Silva, M. (2025). Tokenization, context windows, and the architecture of language models for educators. Computers and Education: Artificial Intelligence, 8, 100245.

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv. https://doi.org/10.48550/arXiv.2307.09288

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., y Polosukhin, I. (2017). Attention is all you need. En Advances in Neural Information Processing Systems, 30 (pp. 5998-6008). Curran Associates.

Veronelli, G. (2015). Sobre la colonialidad del lenguaje. Universitas Humanística, 81, 33-58. https://doi.org/10.11144/Javeriana.uh81.scl DOI: https://doi.org/10.11144/Javeriana.uh81.scdl

Xuan, W., Yang, K., Tan, Y., Tian, X., Qin, Y., Zhang, M., et al. (2025). MMLU-ProX: A multilingual benchmark for advanced large language model evaluation. arXiv. https://doi.org/10.48550/arXiv.2503.10497 DOI: https://doi.org/10.18653/v1/2025.emnlp-main.79

Descargas

Publicado

2026-07-27

Número

Sección

Artículos

Cómo citar

Literacidad en IA: Funcionamiento, Sesgos y Operatividad de los LLMs. (2026). Espila, 8(2), 118-133. https://doi.org/10.61454/yhetdk61

Artículos similares

11-20 de 94

También puede Iniciar una búsqueda de similitud avanzada para este artículo.