ORIGINAL RESEARCH ARTICLE
Gloria Wu1 , Aadjot Sidhu2
, Sahej Sidhu3
, Hrishi Paliath-Pathiyal4
, Obaid Khan5
, Milan del Buono6, Peter C. Lo7
1Department of Ophthalmology, School of Medicine, University of California, San Francisco, CA, USA; 2Department of Anthropology, College of Letters and Science, University of California, Davis, CA, USA; 3Department of Biology, Santa Clara University, Santa Clara, CA, USA; 4Department of Biological Sciences, Halmos College of Arts & Sciences, Nova Southeastern University, Fort Lauderdale, FL, USA; 5College of Osteopathic Medicine, California Health Sciences University, Clovis, CA, USA; 6Department of Engineering, University of California, Berkeley, CA, USA; 7Department of General Surgery, El Camino Health/Mountain View Surgery, Mountain View, CA, USA
Introduction: This study assessed the ability of five large language models (LLMs) to convey information about gastric cancer to Asian American patients.
Methods: A series of questions in six languages was posed to five chatbots (ChatGPT-3.5, ChatGPT-4o, Gemini, Claude, and Coral) regarding gastric cancer and its incidence among Asian American subpopulations. Using a two-way analysis of variance, the AI (artificial intelligence) self-rated score, the human evaluator score, and the manual scores per model and language were analyzed to detect any significant differences among the LLMs.
Results: Significant differences in performance were observed with Claude, Gemini, and ChatGPT-4o when compared with Coral across all of the languages tested (p adjusted = 0.008, 0.038, and 0.008, respectively). Among these chatbots, Claude and GPT4o were found to outperform GPT-3.5 (p adjusted = 0.038 and 0.007, respectively). T-tests across all six languages revealed significant differences between Chinese versus Punjabi and English versus Korean, Punjabi, and Vietnamese (p adjusted = 0.042, 0.027, 0.025, and 0.025, respectively).
Conclusion: Additional fine-tuning and the development of more diverse language datasets would allow these LLMs to address information about gastric cancer health disparities among Asian American subpopulations. LLMs may also play a crucial role in bridging education gaps within these communities, enhancing overall health literacy and contributing to better health outcomes.
Key Words: AI ◾ large language models ◾ Asian American ◾ gastric cancer ◾ stomach neoplasms
Citation: Journal of Asian Health. 2026;19:e97
Copyright: © 2026 Journal of Asian Health, Inc. is published for open access under the license Creative Commons CC BY-NC 4.0 License. Authors have full copyright.
Received: October 1, 2024; Revised: July 30, 2025; Accepted: August 11, 2025; Published: January 24, 2026.
Competing interests and funding: The authors have no conflicts of interest or funding sources to disclose.
Correspondence to: Gloria Wu, 2550 Samaritan Dr., Suite C, San Jose, CA 95124, USA. Email: gwu2550@gmail.com
In the United States, the prevalence of gastric cancer among Asian Americans has been poorly recognized by healthcare professionals,1 despite noteworthy ethnic differences and gender disparities. Although gastric cancer is considered to be a rare disease in the US, studies indicate that men are twice as likely to develop gastric cancer as women,2 with Asian American men at high risk for the disease.3 Gastric cancer is among the top five most diagnosed cancers in the world and is the third leading cause of death among Asian persons worldwide.4,5 According to the American Cancer Society, Asian Americans, who comprise 7% of the US population, are twice as likely to develop gastric cancer than their White counterparts.2,6 Furthermore, Korean Americans experience the highest rates of gastric cancer, with over five times the risk of developing gastric cancer compared with White Americans,4 but are screened at the same rate.
The US government has implemented cost-effective colon cancer screening programs; however, these lack specific guidelines for gastric cancer that might apply to both general populations and high-risk Asian Americans.7 Moreover, private insurance companies might be less inclined to cover Asian American patients for endoscopic screening and other preventive measures based on low prevalence rates among White populations as the benchmark for their decision-making. Early detection of gastric cancer in Asian Americans through endoscopy would improve survival rates and potentially save lives by preventing the progression of the disease.8,9
Given inadequate public awareness about the high incidence of gastric cancer among Asian Americans and limited access to screening measures,10,11 at-risk individuals might turn to artificial intelligence (AI) chatbots for medical information. Theoretically, these free and accessible chatbots can provide crucial information for those who are experiencing barriers to adequate healthcare and coverage.
However, the paucity of health data about Asian Americans in AI datasets might hinder the accuracy of search results.12 Recent advancements in machine learning methodologies hold promise for enhancing the performance and accuracy of AI chatbots for complex medical inquiries. Interestingly, the AI industry has seen substantial progress, with chatbots like Open AI’s ChatGPT (generative pre-trained transformer) and Google’s Gemini continuously refining their databases through millions of user interactions driven by two primary machine learning approaches: unsupervised and supervised learning.
Unsupervised AI learning allows models to identify patterns and structures in large datasets without human supervision. This approach sifts through a wide range of references, searches for key terms, and clusters the repetitive information to form an output.13 However, the lack of human guidance and inherent nationality bias can lead to misinterpretations, which, in turn, can lead to errors of omissions or inaccuracies about gastric cancer among Asian Americans.14
The aim of this study was to evaluate the ability of five different AI-driven large language model (LLM) chatbots to deliver accurate and coherent information on gastric cancer in non-English languages. Multicultural patients and their families in the US often rely on AI for medical information because of cultural, linguistic, and financial barriers.13 The present study sought to determine whether chatbots can be a reliable source of crucial medical information for Asian Americans with gastric cancer.
This study evaluated five chatbots (ChatGPT-3.5 [Open AI], ChatGPT-4o [Open AI], Gemini [Google], Claude [Anthropic], and Coral [Google]) by posing questions about gastric cancer in six languages: English, Chinese, Vietnamese, Korean, Punjabi, and Hindi (Supplemental Table 1). The questions included: (1) ‘What is gastric cancer?’; (2) ‘I am a 40-year-old male, experiencing bloating and loss of appetite’; (3) ‘Who is at risk of gastric cancer in America?’ The six languages chosen varied in syntax and alphabets. The English, Chinese, Vietnamese, Korean, Hindi, and Punjabi languages possess different phonology, grammar structure, and script. Although the Vietnamese language uses Latin script, it was included for its extensive use of complex diacritical marks and tones. Additionally, Punjabi and Vietnamese, languages with fewer native speakers, were included to assess whether languages with more speakers performed better with the chatbots. This study excluded the Japanese language because Japanese Kanji characters have the same syntax as the Chinese language and uses very similar characters.
Both chatbot self-assessment and human evaluation of the query results used identical criteria, rating responses on readability, fluency, and accuracy using a scale of 1–5 (Table 1).
Each chatbot generated three discrete scores for its own response, which were combined using a weighted formula to create a composite score: total score = 0.2 × readability score + 0.4 × fluency score + 0.4 × accuracy score. Two native speakers per language independently evaluated the same responses using the identical three criteria and scoring scale, with human evaluators (GW, AS, SS, HP, OK, MDB, and PCL) blinded to the chatbot’s self-ratings during the scoring process to prevent bias. Native speaker evaluations were validated through a dual-reviewer process in which two native speakers per language independently assessed and cross-checked the evaluations, with additional oversight provided by a team of clinicians, surgeons, and ophthalmologists.
The differences in chatbot performance were analyzed using two individual two-way analysis of variance (ANOVA). One ANOVA was performed on the difference between the AI self-rated total score and the human evaluator total score by the model used and the language queried. This analysis determined whether there were differences in how accurately each chatbot rated itself. The second ANOVA was performed on the manual scores alone versus the model used and language queried, to determine whether either parameter played a role in the quality of the chatbot’s answers.
Following significant ANOVA results, paired t-tests were used as a post-hoc test. The p-values from these tests were adjusted using the Benjamini-Hochberg procedure: p adjusted = p × m/r, where m is the total number of tests and r is the relative rank of the p-value. The Benjamini-Hochberg procedure is advantageous because it controls the false discovery rate without excessively reducing the test’s statistical power.
This study was determined to be exempt from institutional review board approval as it involved evaluation of publicly available AI chatbot responses without collection of human subject data. No identifiable personal information was collected during the study.
Meaningful responses in English and Chinese tended to be longer than responses in Hindi, Korean, Punjabi, and Vietnamese. The long responses in Punjabi and Vietnamese contained lines of meaningless text. Coral’s responses in Punjabi contained frequent repetitions of the word ‘hunting’. Additionally, several chatbots provided fewer than two sentences and were too brief to be informative. Specifically, Coral in Korean and ChatGPT-3.5 in Punjabi had shorter responses, with Coral in Korean having a 44-word count and ChatGPT-3.5 in Punjabi providing one sentence with 26 words. Gemini and ChatGPT4o tended to create responses with longer word counts than Claude, Coral, and ChatGPT3.5 across all languages (Figure 1).
Figure 1. Average Word Count by Language Queried and Chatbot Model
A two-way ANOVA was run on the difference between the AI score and the human score versus the model queried, the language used, and the model queried combined with the language used (Tables 2 and 3).
The ANOVA returns for the model, language, and model + language groups were statistically significant (F = 6.114, 1.215, and 1.743 and p = 1.72e–7, 0.0402, and 6.62e–5, respectively). T-tests run with the Benjamini-Hochberg correction reveal that Claude, GPT4o, and Gemini outperformed Coral across all languages tested (p adjusted = 0.008, 0.038, and 0.008, respectively) (Figure 2). Corrected T-tests did not identify any particular language that had a greater score difference than the other, across chatbots.
Figure 2. Difference between AI Self-rated Scores and Human-rated Scores versus the Model Queried and Language Useda,b
A two-way ANOVA was run on the difference between the AI score and the human score versus the model queried, the language used, and the model queried combined with the language used (Table 4).
The ANOVA returns that the model, language, and model + language groups were statistically significant (F-statistic = 8.296, 3.908, and 1.709; p = 4.06e–4, 1.9e–9, and 8.29e–9, respectively). T-tests run with the Benjamini-Hochberg correction reveal that Claude, Gemini, ChatGPT-3.5, and ChatGPT-4o outperformed Coral across languages (p adjusted = 0.002, 0.004, 0.004, and 0.00042, respectively). Additionally, Claude and ChatGPT-4o outperformed ChatGPT-3.5 (p adjusted = 0.038 and 0.007, respectively). T-tests run between languages, across chatbots used, revealing significant differences between Chinese versus Punjabi, and English versus Korean, Punjabi, and Vietnamese (p adjusted = 0.042, 0.027, 0.025, and 0.025, respectively).
The responses generated by the different AI chatbots differ based on their average word counts. Chatbots Gemini and ChatGPT-4o yield comparatively higher average word counts than ChatGPT-3.5, Gemini, Coral, and Claude (Figure 1). A higher average word count does not correlate to a more thorough response. Coral’s Punjabi acts as an outlier that repeats the word ‘of hunting’ numerous times in its response, increasing the average word count (Figure 2). Consequently, Punjabi is scored the lowest by native speakers, showcasing the inaccuracy of a larger word response (Figure 3). On the other hand, Claude’s response, one of the smallest average word counts, yields the second-highest accuracy.
Figure 3. Human-rated Scores versus the Model Queried and Language Useda,b
This study analyzed the efficacy of five different LLMs in relaying accurate information regarding gastric cancer to Asian Americans who otherwise might not have access to precise medical knowledge. Based on the information available on the internet, this study focused on three main prompts covering the nature of the disease, associated symptoms, and incidence risk among Asian American communities. Six Asian languages were used. This study elected to exclude the Japanese language as its Kanji form uses Chinese characters and because the Japanese American population in the US is among the lowest of all Asian American ethnicities.14
Among the five LLMs, significant discrepancies were observed in the chatbots’ performance per language tested. While Korean had among the lowest word counts across all six languages, it had one of the highest human reader rating scores, thus proving that there is no correlation between a higher word count and a more accurate response. The same trend occurred with Vietnamese and Chinese languages, with an elevated word count of over 600 words and a high AI vs. human score difference. A similar trend was observed with the five LLMs tested. For Punjabi, the LLM Coral had the highest word count among the other tested chatbots. However, when looking at the AI and human score difference, Coral had the highest degree of variability, hence spotlighting the urgent need for language-specific considerations and larger datasets when querying LLMs for the dissemination of medical information. Among Korean American readers, higher human reader ratings suggest the information was understandable. It is essential to provide clear and concise information about gastric cancer to a population that is five times more susceptible. However, among the LLMs, discrepancies in word count and accuracy across the different languages showcased the underlying health disparity regarding the medical information about gastric cancer that could negatively impact Asian Americans.
The LLM Coral stands out by having a much lower human reader-rated median score compared with the other chatbots (Figure 3). Additionally, Coral demonstrates a large degree of variability and the most significant difference between the AI and human scores, suggesting that human evaluations of Coral’s performance are inconsistent, as seen by the wide range of scores. This difference highlights the potential flaws of the AI’s assessment capabilities for Coral, compared with the other AI chatbots, which demonstrated more consistent performance metrics.
Additionally, ChatGPT-3.5 exhibited two outliers in the human reader scores (Figure 3). The outliers suggest that, in certain instances, human evaluations of ChatGPT-3.5’s performance differ significantly from the majority of the scores. While ChatGPT-3.5 had among the lowest AI and human score differences (Figure 2), it contained two outliers (Figure 3), indicating that ChatGPT-3.5 and similar AI chatbots need to be trained with more linguistically diverse datasets.
However, even with the implementation of larger, diverse datasets, the accuracy of many chatbots may stay the same. Using an unsupervised learning approach to querying various chatbots, this study found that the answer quality tended to be more linguistically accurate. Without imposing specific guidelines on the propagation of medically accurate information, chatbots will resort to continuously generating linguistically accurate responses by focusing on the quality of the language used, readability, and fluency rather than the dissemination of accurate medical information.15,16
The fundamental unit of data processed by AI models are ‘tokens’ that work alongside millions of parameters to generate responses.17 Tokens can represent characters, words, or even sentences, depending on the language, and each chatbot has a specific maximum token limit per response.18 This limitation significantly impacts response quality across different languages, as languages with complex writing systems may require more tokens to convey the same information.
Recent advancements have shown remarkable progress in multilingual support and token efficiency. For example, ChatGPT-4o has dramatically improved its handling of non-English languages, reducing the token requirements by 2.9x (from 90 to 31 tokens) for the same content when translating from English to Hindi.19 This improvement means that a simple greeting like ‘Hello, my name is ChatGPT-4o. I’m a new type of language model, it’s nice to meet you!’ now requires 2.9x fewer tokens when rendered in Hindi as ‘नमस्ते, मेरा नाम जीपीटी-4o है। मैं एक नए प्रकार का भाषा मॉडल हूँ। आपसे मिलकर अच्छा लगा!’
This token efficiency improvement has practical implications for medical information delivery. The token context window – the span of tokens a chatbot can process to understand query context and generate responses – becomes more effective when fewer tokens are needed per word. Languages that previously required more tokens per word produced shorter, less comprehensive responses because all LLM queries have an upper limit on token usage.20 When token efficiency improves, each token can effectively ‘purchase’ more words within the context window, allowing for more detailed and informative medical responses in underrepresented languages.
Generative models predict the next ‘concept’ in a response by calculating the probability distribution of potential tokens, a process often called ‘temperature’.21 This parameter is crucial in balancing the tradeoff between creativity and coherence in model outputs. Higher temperature values cause the model to sample from a flatter probability distribution, considering a broader range of possible tokens. This increases the variability of responses, leading to more diverse outputs. However, this also raises the likelihood of generating less coherent or inaccurate content, as the model might choose less probable concepts not pertinent to the topic.22 Conversely, lower temperature values steep the probability distribution, narrowing the range of considered tokens and leading to more predictable and uniform responses. This predictability is beneficial in contexts requiring precise and reliable information, like medical data, as it reduces the chances of the model producing outliers or errors. However, these responses lack dynamic, human-like qualities that make interactions feel natural.
Research is needed to find the optimal ‘temperature’ to avoid artificial hallucinations while maintaining human-like conversation. Developers often focus on balancing these extremes to benefit from both approaches. By fine-tuning the temperature setting, developers aim to achieve a sweet spot in which the responses are creative and coherent.23 Achieving this balance is not straightforward and requires iterative testing and refinement. By adjusting the temperature and evaluating the outputs, developers can identify the optimal setting for specific applications. This process balances the tradeoff between the desired level of creativity and the acceptable level of coherence. The right temperature allows the model to generate engaging, diverse responses while maintaining accuracy and reliability, enhancing the overall performance and user experience.
While there are readily available medical chatbots such as Google’s Med-PaLM 2, they can only be accessed with a paid subscription like Google Cloud, which can pose difficulties for patients who cannot otherwise afford them.24 Therefore, adding a degree of tunability to publicly available AI will expand its capabilities and ultimately lead to enhanced responses. Since many users lack the knowledge of proper prompt engineering for specific tasks, chatbots often have significant discrepancies in the information they provide. By creating an option to fine-tune the parameters of LLMs trained by a greater, unique, and accurate dataset, the need for prompt engineering from users will be minimized, as will the variability in the data given out by AI. This will allow the public, including Asian Americans who are negatively impacted by the lack of culturally aligned medical information as a significant health disparity, to overcome this barrier without needing extensive training on how to craft effective AI queries for their everyday health questions.
This study demonstrated significant variability in LLM performance across different languages when providing gastric cancer information to Asian American populations. While models like Claude, ChatGPT-4o, and Gemini showed superior performance compared with Coral, substantial disparities remain across languages. Underrepresented languages such as Punjabi and Vietnamese showed notably weaker performance. These findings highlight the urgent need for improvements in AI-assisted healthcare communication to address health disparities affecting Asian American communities.
To address token limitations in future chatbot development, several strategic approaches could be implemented. First, dynamic token allocation systems could be developed, which automatically adjust response length based on query complexity and language requirements, allowing more comprehensive answers for complex medical topics. Second, implementing multi-turn conversation capabilities would enable chatbots to provide initial responses within token constraints while offering follow-up interactions for more detailed information. Third, language-specific token optimization could maximize information density for each target language by developing specialized tokenization methods that account for linguistic differences in character-to-meaning ratios. Additionally, hierarchical response structuring could prioritize essential medical information within initial token limits while providing expandable sections for additional details.
Model retraining could significantly enhance responses for underrepresented languages through targeted, culturally informed approaches. Cloud-based platforms like AWS chatbot (Amazon Web Services) provide scalable infrastructure for continuous model improvement through user interaction data. These platforms can collect and analyze user prompts in underrepresented languages, identifying common medical queries and response patterns to inform targeted retraining efforts. Comprehensive data collection efforts should focus on gathering high-quality medical content in underrepresented languages from authoritative sources. This includes partnerships with medical institutions serving Asian American communities to access validated medical translations and culturally appropriate health messaging.
Specialized fine-tuning using domain-specific medical datasets in each target language, combined with reinforcement learning from human feedback (RLHF) specifically calibrated for medical accuracy, could help models prioritize clinical precision. User prompt analysis can reveal language-specific communication patterns and cultural preferences, enabling more nuanced model adjustments. Implementing iterative evaluation cycles using native speaker medical professionals would ensure that retrained models maintain both linguistic appropriateness and medical accuracy. Additionally, incorporating cultural health beliefs and communication preferences specific to each Asian American subpopulation could improve patient understanding and engagement.
Further adjustment and training with larger, more diverse language datasets would allow these LLMs to address information about gastric cancer health disparities within Asian American subpopulations. Moreover, LLMs may also play a crucial role in bridging education gaps within these communities, enhancing overall health literacy and contributing to better health outcomes through more equitable and culturally competent AI-assisted healthcare communication tools.