用于晚期头颈部恶性肿瘤管理的大型语言模型的可靠性： ChatGPT 4 和 Gemini Advanced 之间的比较。Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced.-医云文献数字医云科研云海量医学决策数据服务

Abstract：

OBJECTIVE: This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI\'s ChatGPT 4 and Google\'s Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes.
METHODS: This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test.
RESULTS: Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2-4]) compared to Gemini Advanced (median 2 [2-3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection.
CONCLUSIONS: While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.

摘要：

目的：本研究评估了两种高级大型语言模型（LLM）的功效，OpenAI的ChatGPT4和Google的双子座高级，为头颈部肿瘤病例提供治疗建议。目的是评估其在支持多学科肿瘤评估和决策过程中的效用。
方法：此比较分析检查了ChatGPT4和Gemini对5例假设的头颈部癌的反应，每个代表不同的解剖亚位点。根据最新的国家综合癌症网络（NCCN）指南，通过两个盲板使用总分歧评分（TDS）和人工智能性能仪器（AIPI）对响应进行了评估。使用Wilcoxon符号秩检验和Friedman检验进行统计评估。
结果：在遵守指南和综合治疗计划方面，两个LLM都提出了ChatGPT4的相关治疗建议，通常优于GeminiAdvanced。ChatGPT4与Gemini高级（中位数2[2-3]）相比，AIPI得分更高（中位数3[2-4]），表明更好的整体性能。值得注意的是,在诱导化疗和手术决策的管理中观察到不一致，如颈部解剖。
结论：虽然这两个LLM都证明了在头颈部肿瘤学的多学科管理方面有帮助的潜力，某些关键领域的差异突出了进一步完善的必要性。该研究支持AI在增强临床决策中的作用，但也强调了不断更新和验证当前临床标准的必要性，以将AI完全整合到医疗保健实践中。