Fiabilidad diagnóstica de la inteligencia artificial multimodal vs. Juicio clínico en la clasificación de fracturas AO/OTA: un estudio exploratorio
Diagnostic reliability of multimodal artificial intelligence vs. Clinical judgment in AO/OTA fracture classification: an exploratory study

Francisco León, Krystell Picón

Resumen


RESUMEN Introducción: La clasificación AO/OTA es el estándar global para la categorización de fracturas, aunque presenta variabilidad interobservador, especialmente en personal en formación. Los modelos de lenguaje de gran escala (LLM) multimodales emergen como una alternativa diagnóstica frente a las redes neuronales convolucionales (CNN) tradicionales. Materiales y Métodos: Estudio de corte transversal y comparativo. Se seleccionaron 50 radiografías de huesos largos, pelvis y superficies articulares. La muestra fue evaluada por el modelo Gemini 3 Flash y por residentes de la UDAOT-ULA, utilizando como Gold Standard el consenso de especialistas. Se empleó el software SPSS v.29 para calcular precisión, sensibilidad, especificidad y concordancia mediante el Índice de Kappa de Cohen (p < 0,05). Resultados: La IA alcanzó una precisión global del 88% y una sensibilidad del 95,7%, superando a los residentes de tercer año (R3: 84%). El análisis de concordancia mostró un índice Kappa de 0,79 para la IA (Sustancial), equiparable a los residentes de cuarto año (0,86) y significativamente superior a los R1 (0,08). No obstante, la IA presentó errores cualitativos graves en regiones de alta complejidad anatómica como la pelvis. Conclusiones: La IA multimodal demuestra ser una herramienta de soporte eficaz y un nivelador de competencias para residentes de niveles iniciales. Sin embargo, debido a limitaciones en el razonamiento espacial complejo, su uso debe integrarse como un sistema de doble lectura bajo supervisión del especialista.
ABSTRACT
Introduction: The AO/OTA classification is the global standard for fracture categorization, although it presents interobserver variability, especially among personnel in training. Multimodal large language models (LLMs) are emerging as a diagnostic alternative to traditional convolutional neural networks (CNNs). Materials and Methods: A cross-sectional comparative study. Fifty radiographs of long bones, pelvis, and articular surfaces were selected. The sample was evaluated by the Gemini 3 Flash model and by residents of the UDAOT-ULA, using a consensus of specialists as the Gold Standard. SPSS v.29 software was used to calculate accuracy, sensitivity, specificity, and concordance using Cohen's Kappa index (p < 0.05). Results: The AI achieved an overall accuracy of 88% and a sensitivity of 95.7%, outperforming third-year residents (R3: 84%). Concordance analysis showed a Kappa index of 0.79 for the AI (Substantial), comparable to fourth-year residents (0.86) and significantly higher than first-year residents (R1: 0.08). However, the AI presented severe qualitative errors in regions of high anatomical complexity such as the pelvis. Conclusions: Multimodal AI proves to be an effective support tool and a competence leveler for junior residents. However, due to limitations in complex spatial reasoning, its use must be integrated as a double-reading system under specialist supervision.

Recibido: 20-06-2026
Aceptado: 31-07-2026
Publicado: 25-09-2026


Palabras clave


inteligencia artificial multimodal; clasificación AO/OTA; fracturas; residentes; fiabilidad diagnóstica; multimodal artificial intelligence; AO/OTA classification; fractures; residents; diagnostic reliability.

Texto completo:

PDF

Referencias


Audigé, L., Bhandari, M., Hanson, B., & Kellam, J. (2005). A concept for the validation of fracture classifications. Clinical Orthopaedics and Related Research, 430, 182–190. https://doi.org/10.1097/01.bot.0000155310.04886.37

Badgeley, M. A., Zech, J. R., Oakden-Rayner, L., Glicksberg, B. S., Liu, M., Gale, W., McConnell, M. V., Percha, B., Snyder, T. M., & Dudley, J. T. (2019). Deep learning predicts hip fracture using confounding patient and healthcare variables. NPJ Digital Medicine, 2, 31, 1-10. https://doi.org/10.1038/s41746-019-0105-1

Chen, K., Stotter, C., Klestil, T., & Nehrer, S. (2022). Artificial intelligence in orthopedic radiography analysis: A narrative review. Diagnostics, 12(9), 2235. https://doi.org/10.3390/diagnostics12092235

Kooistra, B. W., Dijkstra, S., Diercks, R. L., & Goslings, J. C. (2011). Reliability of the AO classification for distal radius fractures. Journal of Orthopaedic Trauma, 25(7), 409–415. https://doi.org/10.1097/BOT.0b013e3181f21503

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310

Langerhuizen, D. W. G., Janssen, S. J., Mallee, W. H., van den Bekerom, M. P. J., Doornberg, J. N., Ring, D., & Schupte, M. J. (2019). What are the applications and limitations of artificial intelligence for fracture detection and classification in orthopaedic trauma imaging? A systematic review. Clinical Orthopaedics and Related Research, 477(11), 2482–2491. https://doi.org/10.1097/CORR.0000000000000848

Langerhuizen, D. W. G., Schupte, M. J., & Doornberg, J. N. (2020). Is deep learning as good as the orthopedic surgeon in classifying distal radius fractures? Archives of Orthopaedic and Trauma Surgery, 140(12), 2021–2027. https://doi.org/10.1007/s00402-020-03539-4

Meinberg, E. G., Agel, J., Roberts, C. S., Karam, M. D., & Kellam, J. F. (2018). Fracture and dislocation classification compendium: 2018. Journal of Orthopaedic Trauma, 32(Suppl 1), S1–S170. https://doi.org/10.1097/BOT.0000000000001063

Moor, C., Banerjee, O., Abad, Z. S., Kulkarni, N., Reyes, N., Singh, P., Zhang, Y., & Shetty, S. (2024). Med-Gemini: High-capability multimodal models for medicine (arXiv:2404.18416). arXiv. https://doi.org/10.48550/arXiv.2404.18416

Müller, M. E., Nazarian, S., Koch, P., & Schatzker, J. (1990). The comprehensive classification of fractures of long bones. Springer-Verlag. https://doi.org/10.1007/978-3-642-75435-7

Olczak, J., Fahlberg, N., Behnam, S., Forghani, B., Kiesel B., Wedin, R., & Gordon, M. (2017). Deep learning, a form of artificial intelligence, open a new window for analysis of musculoskeletal radiographs. Acta Orthopaedica, 88(6), 581–586. https://doi.org/10.1080/17453674.2017.1368812

Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tan, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamwa, P., Schärli, N., Fisher, A., Bebensee, A., Drucker, J., Schupmann, C., Cole, A., ... Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2

Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Ting, L., & Carin, L. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930–1940. https://doi.org/10.1038/s41591-023-02448-8

Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7


Enlaces refback

  • No hay ningún enlace refback.


 

Depósito Legal Electrónico: ME2016000090
ISSN Electrónico: 2610-797X

DOI: https://doi.org/10.53766/GICOS

Se encuentra actualmente registrada y aceptada en las siguientes base de datos, directorios e índices: 

 
 
   
    
    
   

Creative Commons License
Todos los documentos publicados en esta revista se distribuyen bajo una
Licencia Creative Commons Atribución -No Comercial- Compartir Igual 4.0 Internacional.
Por lo que el envío, procesamiento y publicación de artículos en la revista es totalmente gratuito.