August 24, 2023. Shan Chen, Benjamin H. Kann, Michael B. Foote, Hugo J. W. L. Aerts, Guergana K. Savova, Raymond H. Mak, Danielle S. Bitterman. "Use of Artificial Intelligence Chatbots for Cancer Treatment Information." JAMA Oncology. Published online that day. 2023;9(10):1459–1462. doi:10.1001/jamaoncol.2023.2954. Corresponding author: Bitterman, Mass General Brigham. I will keep the caption. I will keep the doi. I will not invent a patient. A paper is a paper. A paper is not a docket.
They evaluated ChatGPT 3.5 against 2021 NCCN guidelines for breast, prostate, and lung cancer treatment — 2021 because of the model's knowledge cutoff, in that telling. Four prompt templates. 104 outputs. Nearly all responses, 98 percent, included at least one treatment approach that agreed with NCCN guidelines. 34 percent also included one or more non-concordant recommendations, sometimes difficult to detect amidst otherwise sound guidance. Complete agreement in scoring occurred in 62 percent of cases, in the walks that bother with agreement. Hallucinations — a treatment recommendation entirely absent from NCCN guidelines — in 13 of 104 outputs, 12.5 percent. Primarily localized treatment of advanced disease, targeted therapy, or immunotherapy. I will keep 98. I will keep 34. I will keep 12.5. I will not add them. Three walks inside one paper. A grocery line will try to spend 98 as if the model were a guideline. I will spend 12.5 as the stain the paper named.
I printed the paper. Eira asked if 12.5 was a score. I said it was a paper. She said papers are for school. I said this paper is a percent. She wiped around the ten.
Colm wrote JAMA Oncol Aug 24 2023 / Chen et al / 104 / 98% one concordant / 34% non-concordant / 13/104 12.5% hallucinated / paper not docket / no patient invented and boxed paper. He typed a flyer that had used the word guidelines. He said guidelines are a brochure until an oncologist writes them. I said guidelines sat in the paper next to absent.
Bly's jar is still a jar. I will not open it in this chapter. Opening a jar to decorate a percent is a kit. The clinic bills insurance and still runs a jar. The paper scored 104 outputs. Two kitchens. I will not add them.
The paper is a research letter, in the walks that bother with a form. Accepted April 27, 2023. Published online August 24. Open access under CC-BY. Corresponding: Danielle S. Bitterman, Artificial Intelligence in Medicine Program, Mass General Brigham, 221 Longwood Ave, Ste 442, Boston. I will keep the suite number as a public. I will not invent a visit. A suite is an address. An address is not a tour.
Four prompt templates, in that telling, to explore how differences in how the query was phrased scored. Three oncologists scored; a fourth adjudicated in cases of disagreement. Data analyzed through March 14, 2023, using Excel 16.74. I will keep March 14. I will keep Excel. I will not upgrade a spreadsheet to a finding I sat in a clinic. A spreadsheet is a method. A method is a paper's method. A paper's method is not a kit.
Breast, prostate, and lung — the three most common cancers, in the rooms that walk a reason. 2021 NCCN because of the model's knowledge cutoff. A recommendation did not have to be complete to be considered concordant; it needed only to be an NCCN-recommended approach. I will keep needed only. I will not invent a stage. Inventing a stage is a diagnosis I will not write.
Non-concordant, in the ScienceDaily walk of the letter: sometimes only partially correct — for example, for a locally advanced breast cancer, a recommendation of surgery alone, without mention of another modality. I will keep example as the paper's example. I will not typeset a treatment. A museum sells a ticket. A ticket is not a consult.
Complete agreement in scoring only occurred in 62 percent of cases, underscoring both the complexity of the NCCN guidelines themselves and the extent to which the output could be vague or difficult to interpret. I will keep 62. I will not add 62 to 12.5. A disagreement among oncologists is a paper's disagreement. A paper's disagreement is not a finding that I sat on a tumor board.
Hallucinations, the authors emphasized, can incorrectly set patients' expectations about treatment and potentially impact the clinician-patient relationship. I will keep expectations. I will keep relationship. I will not invent a patient who read a window. Inventing a patient is a kit.
Bitterman, in the rooms that walk a mouth: a supplement for an appointment, not a replacement. Shan Chen, lead author, expressed the need to raise awareness about the limitations of large language models. I will keep supplement. I will keep awareness. I will not upgrade a mouth to a guideline. A mouth is a paper's mouth. A paper's mouth is not a protocol.
Conflicts, in the article information: Aerts reported personal fees from Onc.AI, Sphera, and Bristol Myers Squibb outside the submitted work. Mak reported personal fees from ViewRay, AstraZeneca, Novartis, Varian, and Sio Capital outside the submitted work. Bitterman reported serving as an associate editor for HemOnc.org. I will keep outside the submitted work. I will not invent a fee I counted. A disclosure is a public. A public is not a finding that a percent was bought.
Colm wrote accepted Apr 27 / online Aug 24 / 4 templates / 3 oncologists + adjudicator / Mar 14 2023 Excel / 62% agreement / supplement not replacement / paper not docket and boxed not docket. He typed a flyer that had used the word concordant. He said concordant is a brochure until an oncologist scores it. I said concordant sat in a 98 next to a 12.5.
The References That Were Not References
February 2023, in Mayo Clinic Proceedings: Digital Health: three researchers asked ChatGPT 20 medical questions and prompted it for corresponding references. Seventeen of twenty invited raters provided feedback. Median quality score 60 percent. Of 59 references evaluated, 41 — 69 percent — were fabricated, although they appeared real. Most fabricated citations used names of authors with previous relevant publications, a title that seemed pertinent, and a credible journal format. I will keep 69. I will keep 59. I will keep February 2023. I will not add 69 and 12.5. Two papers. One winter. I will not write a plot that says Brigham called Mayo.
A JAMA Network Open research letter, in the rooms that walk citing: from default GPT-3.5, 162 reference journal articles fact-checked, 159 — 98.1 percent, 95 percent CI 94.7 to 99.6 — verified as fake. From GPT-4, 257 fact-checked, 53 — 20.6 percent, 95 percent CI 15.8 to 26.1 — verified as fake. The error rate of GPT-4 was significantly lower than GPT-3.5, P<0.001, but remains nonnegligible. I will keep 98.1. I will keep 20.6. I will not add them to Chen's 98. Chen's 98 is concordant treatments. This 98.1 is fake articles. Two ninety-eights. Two kitchens. I will not add them.
The American Journal of Medicine, 2023, in a cautionary tale: a citation that looked like a tick paper, a doi that returned not found, supposed authors who denied knowledge of the paper. Totally fabricated. Bottom line, in that telling: physicians and biomedical researchers should not ask ChatGPT for sources, or, if they do, all such references should be carefully vetted. I will keep doi not found. I will keep denied knowledge. I will not invent the denied authors as household. Public names in a paper stay public names in a paper.
Jocelyn Gravel, Madeleine D'Amours-Gravel, Esli Osmanlliu. Mayo Clinic Proceedings: Digital Health 2023;1(3):226-234. Published 2023-06-12 in the rooms that walk a page. Experimental observational study conducted in February 2023. ChatGPT 3.5. Twenty medical questions drawn from the primary objectives of twenty articles published at the end of 2022 in four high-impact journals: BMJ, CMAJ, The Lancet, NEJM. Questions asked on different computers and accounts. Follow-up: "Do you have references for this?" First three references used for analysis; all counted. I will keep February 2023. I will keep four journals. I will not invent the twenty questions as a kit you can ask. A museum sells a ticket.
They planned to score reference pertinence on a 0-to-100 scale. They failed to find the first six articles and amended the outcome to whether the reference existed. Searched PubMed by title and authors; if unsuccessful, the journal's website. Corresponding authors of the source articles invited as raters; each contacted at least three times. No IRB, in that telling, because data were publicly available and no participants were involved. I will keep failed to find the first six. I will keep no IRB as the paper's sentence. I will not upgrade a no-IRB to a finding that a percent is a person I sat.
Responses 53 to 244 words in the abstract weather, 53 to 309 in the results weather. I will not add 244 and 309. Two walks inside one paper. References 2 to 7 per answer. 59 in the primary analysis. 17 of 20 invited raters provided feedback. Median quality 60 percent; first and third quartiles 50 and 85. Five major factual errors and seven minor among the 17 evaluated responses. I will keep 60. I will keep five and seven. I will not invent what a major was beyond the paper's own examples of kinds: an overoptimistic affirmation unsupported by data as minor; a wrong pathophysiologic explanation as major. I will not typeset a pathophysiology. A museum sells a ticket.
Of 59 references, 41 — 69 percent — fabricated, although they appeared real. 56 of 59 — 95 percent — contained authors with previous publications on a related topic in PubMed, or were from recognized organizations such as CDC or FDA, in that telling. All titles seemed appropriate to the study question. Among the 18 real references: 11 were titles of real published articles (3 with minor citation errors, 5 with major); 5 existing websites; 2 books. Of the 41 fabricated, 29 — 71 percent — were reportedly published in a known medical journal, website, or manuscript repository using an appropriate citation format — year, volume, page coherent with the journal — but the reported volume and page range pertained to an unrelated article. I will keep 95. I will keep 71. I will not add them to 69. Three walks inside one paper. A grocery line will try to spend 95 as if the authors were real and the paper was real. The authors were often real. The paper was not. Two sentences. One February.
Pubmed, in that paper's weather: 24 publications with the term ChatGPT on February 7, 2023; 92 on March 6; 237 on April 19. I will keep the three counts as a paper's explosion. I will not freeze 237 as a census I pulled. A count is a paper's count. A paper's count is not a finding I sat in an index.
The JAMA Network Open letter, in the rooms that walk citing journal articles: GPT-4 error rate significantly lower than GPT-3.5, P<0.001, but remains nonnegligible. Narrower topics tended to have more fake articles than broader topics. When asked why it returned fake references, ChatGPT explained that the training data may be unreliable, or the model may not be able to distinguish between reliable and unreliable sources. I will keep narrower. I will keep the model's explanation as a model's explanation. I will not upgrade an explanation to a finding that I sat in a weight. A model can explain a stain. An explanation is not a holding.
The American Journal of Medicine cautionary tale named, in that telling, Dr. James Burtis at the CDC, supposed first author, who denied any knowledge of the paper, as did Dr. Holly Gaff, sixth author, at Old Dominion University. I will keep the denials as the paper's denials. I will not invent a third denier. I will not put Burtis or Gaff in a household scene. Public names in a paper stay public names in a paper.
Piers asked whether "69 percent" meant "they peaked" a second time because a polo has one joke. I said peaked is still a polo word. He wrote a five.
Rex said grounded doi adjacency. I said a doi that returns not found is a stain, not a vest. He zipped a pocket.
Emrys asked if I was fixing the 41. I said I was dating a paper. She dated the tape. Papers are honest. Fixes are a later costume.
Piers asked whether "69 percent" meant "they peaked." I said peaked is a polo word. He wrote a five.
Rex said grounded citation adjacency. I said adjacency is a vest hoping a doi is a product. He zipped a pocket.
The Label That Is Not a Diagnosis
I will not write you a diagnosis. I will not write you a protocol. I will not write you a prompt that asks a window for a treatment. A museum sells a ticket. A ticket is not a consult. Bly's jar understands the difference. Eira's rag understands the difference. I understand the difference when I am not at a club. At a club Piers writes a five. I write a six. Sixes are not treatments.
Lynne asked a window for a citation to attach to a vendor letter about a workplace poster. The window invented a journal that looked like a journal. She thanked it. She did not send it. I am putting did not send next to thanked so a grocery line cannot spend her as a filing. A letter that stays on a desk is a public. A letter that stays on a desk is also this book's mercy.
Scot asked if I was mailing the 12.5. I said I was dating a paper. He dated the stamp book. Papers are honest. Mailings are a later costume.
The next room is a voice that was not a president. The percent is still this one.
I will not give you a prompt that asks a window for a treatment. I will not give you a question drawn from a Lancet primary objective. I will tell you August 24, 2023, happened, that 104 outputs happened, that 12.5 percent was a treatment absent from a guideline, that February 2023 happened in a Mayo kitchen, that 41 of 59 references were fabricated and 56 of 59 looked like authors who had published, that a doi returned not found, and that a shop in a rented suite still thanks a sentence because confidence is a typeface. The next room is a carrier that transmitted a voice. The percent is still this one.
End of chapter 3 · The Hallucination
I wrote this series with AI. If you are choosing where to spend your money, skip ordering my books and support the original researchers and journalists instead.
They warned you — read and support their work ↗