Large Language Models in Clinical Neuropsychology

Large Language Models in Clinical Neuropsychology

Authors: Matt Hutnyan, Cynthia Beaulieu, Amanda Ball

Synopsis: Electronic methods of writing have existed for decades. In today’s face-paced digital age, electronic platforms using artificial intelligence (AI) methods have become sufficiently sophisticated that concise, coherent, and meaningful write ups can be created from collation of large amounts of information generated from informative prompts and instructions (generative AI). This includes their use in creating professional reports, presentations, and social media content. This Tech Tip provides a list of 4 common generative AI platforms that are used to generate content, their strengths and weaknesses, who the platform is ideally suited for, and whether or not any of the platforms are HIPAA compliant for use with patient health information (PHI).

AS A REMINDER: Regardless of why you are using a generative AI platform, always, always, always remember the “human-loop” rule, which is the practice of you, the writer, ensuring you are the final authority on what you write and what you submit. Review, review, review any written product created with any form of assistive technology, including not just generative AI platforms but also the more common voice-to-text options available widely.

Basic Definition of Large Language Models (LLM): Advanced computer programs that read vast amounts of human generated text to learn how words are naturally stringed together (i.e., natural language process; NLP). By understanding which words usually follow one another,  LLM can write original text, answer questions, summarize long documents, and explain complex ideas

Advanced Definition of Large Language Models: Neural network-based systems evaluate the contextual relationship between words and then predict (anticipate) the most logical next progression of a statement. These systems are trained on large quantities of text to learn probability distributions over token sequences (individual words, statements, phrases). They generate text by predicting the most likely next token given the prior context. This allows generation of summaries, responses to questions, and explanations.

Examples of platforms that use LLMs:

  • ChatGPT
  • Claude
  • Microsoft Co-Pilot
  • Gemini
  • Grok
  • BastionGPT
  • MedLM
  • Med-PaLM
  • OpenScholar
  • OpenEvidence
  • Perplexity

Strengths of LLMs:

  • Can handle many language tasks without task-specific programming.
  • Good at generating fluent, human-like text.
  • Can generalize across topics and domains.

Limitations of LLMs:

  • May generate incorrect information confidently (i.e., “hallucinations”).
  • Do not truly “understand” concepts in the human sense; they model patterns in data.
  • Can inherit biases from the data they are trained on.
  • Performance depends heavily on the quality of the prompt and training data.

Applications in Clinical Neuropsychology:

Overall Description (see Table): Remember the adage, “garbage in, garbage out.” The quality of AI-generated output is highly dependent on the quality, specificity, and context of the information provided in the prompt (see example below on how to compose an effective AI prompt).

PLATFORMCREATORUSER OPTIONSREVIEWSBEST SUITED FORMORE INFORMATION
ChatGPTOpenAI

Free version available

Fee for elevated version (Plus version)

HIPAA-Ready version:
ChatGPT for Healthcare
(no training on data)

PROs:
Most popular
Active user base
Specialized tools
Versatility

CONs:
Message limitations

Privacy concerns for use in organizations

Reports of “repetitive” creation

Assistance with writing, some research, vibe coding, live conversations, creation of media and presentations

openai.com/chatgpt

Top 15 AI Platforms in 2026 (Tested & Ranked)

I tried 70+ best AI tools in 2026 | TechRadar

ClaudeAnthropic

Free version with limited usage

Fees for elevated versions (Pro, Team, Enterprise)

HIPAA-Ready version:
Claude for Healthcare
(no training on data)

PROs:
Few hallucinations
Excellent reasoning
Best for coding
Long document analysis
Cleanest output
Workflow enablement
Less AI-like output

CONs:
Limited image generation
Overly cautious
Limited mobile capability

Assistance with long documents, coding

Best for developers, researchers, analysts.

claude.ai

Top 15 AI Platforms in 2026 (Tested & Ranked)

The 18 Best AI Platforms in 2026 – Tested & Reviewed | Lindy

GeminiGoogle

Free version available

Fee for elevated version (Gemini Advanced)

HIPAA-Ready version:
Gemini for Google Workspace
(no training on data)

PROs:
Multimodal visual and voice response times

Auto-verification and fact checking

Feels like an assistant

CONs:
Creativity/brainstorming and voice-tone control

May lean too heavily on past searches

Individuals who want smooth integration with Google Workspace and across personal Google devices, including Gmail, Docs, Sheets, Slides, Drive, Calendar

gemini.google.com

I tried 70+ best AI tools in 2026 | TechRadar

CoPilotMicrosoft

Basic version in Bing and Windows

Fee for elevated versions (Pro, MS 365)

HIPAA-Ready version:
Copilot for Microsoft 365
(no training on data)

PROs:
Best integration with office products

Enterprise-grade security and compliance

CONs:
Requires MS 365
Steep learning curve
Chat experience lagging behind other platforms

Individuals who want smooth integration with Microsoft products, including Word, Excel, PowerPoint, Outlook, Teams, OneNote

copilot.microsoft.com

Top 15 AI Platforms in 2026 (Tested & Ranked)

PerplexityPerplexity AIFree version available.
Fee for elevated version (Pro).
Enterprise options available. Compliance and privacy features should be verified directly with the vendor.

PROs: Excellent for literature searches and background research. Provides citations and source links. Strong web search and information synthesis capabilities. User-friendly interface.

CONs: Quality depends on source material. May generate inaccurate or incomplete information. Less suited for long-form writing and content creation than some other AI tools.

Literature reviews, background research, source identification, fact checking, and rapid synthesis of information from multiple sources.

perplexity.ai

Best AI Research Tools 2026: Elicit, Perplexity & NotebookLM

OpenEvidenceOpenEvidenceRequires account created with healthcare license (NPI must be submitted)

PROs:
Specific medical search engine. Highest performer on four medical benchmarks when compared to general purpose or foundation models. Official AI partnership with leading medical journals (NEJM, JAMA, Nature, Cochrane, NCCN).

CONs:
Primarily US focus. Response engine rather than comprehensive clinical tool. Limited transparency on how queries are processed. Requires account created with healthcare license (NPI must be submitted).

Literature reviews, evidence-based practice.

About | OpenEvidence

https://doi.org/10.1002/jac5.70237

Example: What makes a good generative AI prompt?

Taking time to think through what you are trying to accomplish is as important as reviewing what is generated from what you ask. The more detailed information you give in the prompt to the AI platform, the better it can respond and ensure it meets your need.

  1. Role of the person. Identify and describe who is doing the asking, which then sets the perspective of the response.
    Bad Prompt: “Give me the best outline for a presentation on DoC.”
    Good Prompt: “I am a clinical neuropsychologist who has 10 years experience assessing and treating brain disorders at an academic medical center in the department of neurology.”
  2. Provide context and background. Include the “why” behind your prompt to prevent assumptions on the part of the AI platform.
    Bad Prompt: “I need to provide the audience with updated information about DoC.”
    Good Prompt: “I will be presenting to a professional group as part of a continuing education course. The audience will be clinical neuropsychologists who have a basic understanding of DoC and can already assess differential diagnostic states of DoC. I need to create a presentation for licensed neuropsychological practitioners who are post-doctorate in their training. The information provided to them needs to be up-to-date and meet guidelines for post-doctoral continuing education.”
  3. Detail the task and specify instructions. Use concise, action-oriented verbs to ensure the AI platform understands the “ask” and can accurately execute the operations needed to answer the request.
    Bad Prompt: “List an outline of subtopics to present on.”
    Good Prompt: “Create an outline for the specific topic “Disorders of Consciousness (DoC): Updated Practice Guidelines”. The presentation will be 45 minutes in duration with 15 minutes for Q&A. Select 3-5 subtopics from the outline, provide bullets for each subtopic, and provide creative take-away statements to promote retention of the key bullets.”
  4. Identify any conditions, constraints, or parameters. Set boundaries on length and tone, and identify priorities to include or exclude.
    Bad Prompt: “Generate a list of 10 questions to include in a post-test.”
    Good Prompt: “Create a 10-question post-test with 8 of the questions in multiple choice format with one correct answer and 2 questions in true/false format. Questions should be based on the outline content and satisfy the 3 learning objectives. Indicate the correct response for each question.”
  5. Specify output requirements and formatting. Detail the structure to ensure the final product is immediately usable and/or available for copy/paste into another application.
    Bad Prompt: “Write a professional title.”
    Good Prompt: “Based on the content outline, target audience, and learning objectives, suggest an appropriate academic but also attention-getting title for the presentation.”

Clinical

  • Patient communication and psychoeducation (e.g., formulation of recommendations)
  • Data extraction (e.g., medical records, natural communication)
  • Testing support (e.g., automation of scoring, generation of psychometric items)
  • Clinical documentation and report writing (e.g., summarization of clinical data)
  • Clinical decision making support (e.g., generation of differential diagnoses)
  • Ai scribes for clinical interviews (e.g. generation of intake note and summary)

Research

  • Data processing and extraction (e.g., cleaning data)
  • Literature synthesis (e.g., preliminary reviews of literature)
  • Hypothesis generation
  • Advanced data analysis (e.g., multimodal data integration)

Education and Training

  • Subject matter tutoring (e.g., exam preparation)
  • Interactive simulations (e.g., preparation for patient interactions)
  • Supervised clinical reasoning practice
  • Supervision augmentation

Peer-Reviewed Articles:

Chlasta, K., Struzik, P., & Wójcik, G. M. (2025). Enhancing dementia and cognitive decline detection with large language models and speech representation learning. Frontiers in Neuroinformatics, 19.

Jaworski III, M., Balconi, J., Santivasci, C., & Calamia, M. (2026). Feasibility of AI-powered assessment scoring: Can large language models replace human raters? The Clinical Neuropsychologist, 40(3), 816-829.

Jin, Y., Liu, J., Li, P., Wang, B., Yan, Y., Zhang, H., … & Wang, Y. (2025). The applications of large language models in mental health: Scoping review. Journal of Medical Internet Research, 27(1).

Kronenberger, O. R., Gottlieb, M. C., & Cullum, C. M. (2026). Large language models in neuropsychology: Emerging applications and ethical considerations. The Clinical Neuropsychologist, 40(3), 753-775.

Kronenberger, O. R., Hutnyan, M., Kaser, A. N., Bullinger, L., Lacritz, L. H., & Cullum, C. M. (2026). Comparing the neuropsychology knowledge base of publicly available large language models. Archives of Clinical Neuropsychology, 41(3).

Moura, L., Jones, D. T., Sheikh, I. S., Murphy, S., Kalfin, M., Kummer, B. R., … & Patel, A. D. (2024). Implications of large language models for quality and efficiency of neurologic care: Emerging issues in neurology. Neurology, 102(11).

Obradovich, N., Khalsa, S. S., Khan, W. U., Suh, J., Perlis, R. H., Ajilore, O., & Paulus, M. P. (2024). Opportunities and risks of large language models in psychiatry. NPP—Digital Psychiatry and Neuroscience, 2(1).

Romano, M. F., Shih, L. C., Paschalidis, I. C., Au, R., & Kolachalama, V. B. (2023). Large language models in neurology research and future practice. Neurology, 101(23).

Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930-1940.

Wolff, B. (2026). Artificial intelligence and natural language processing in modern clinical neuropsychology: A narrative review. The Clinical Neuropsychologist, 40(3), 728-752.

Additional Resources:

What Are Large Language Models (LLMs)? | IBM

AI Demystified: Introduction to large language models | University IT

AI vs Human Thinking: How Large Language Models Really Work

One Useful Thing

What Are AI Hallucinations? | IBM

OpenAI Prompt Engineering Guide
 https://platform.openai.com/docs/guides/prompt-engineering

Microsoft Copilot Prompting Basics
https://support.microsoft.com/en-us/topic/get-started-writing-prompts-in-microsoft-copilot-86ccf1a2-d4b4-4b59-bc20-88f7c7044ddb

Microsoft Copilot Prompt Gallery
 https://adoption.microsoft.com/en-us/copilot/prompt-gallery/

NIST AI Risk Management Framework (AI RMF)
 https://www.nist.gov/itl/ai-risk-management-framework

NIST Generative AI Profile
https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

Share Your Experience

Have you used generative AI tools in your clinical, research, educational, or administrative work? We welcome feedback, suggestions for future tech tips, and examples of practical applications. Please reach out to Amanda Ball at ballama@ohsu.edu

Dear AACN members, please log in to share your comments or questions here.

Disclosure and Disclaimer: Products and services mentioned in this resource are provided as examples for informational purposes only and do not constitute endorsement by the AACN DTI Committee or AACN. Committee members report no relevant financial relationships with the companies discussed. Because product features and compliance information may change, readers should independently verify current specifications and determine suitability for their practice setting.