Benchmarking Multiple Large Language Models for Automated Clinical Trial Data Extraction in Aging Research

Large-language models (LLMs) show promise for automating evidence synthesis, yet head-to-head evaluations remain scarce. We benchmarked five state-of-the-art LLMs—openai/o1-mini, x-ai/grok-2-1212, meta-llama/Llama-3.3-70B-Instruct, google/Gemini-Flash-1.5-8B, and deepseek/DeepSeek-R1-70B-Distill—on...

Full description

Saved in:
Bibliographic Details
Main Authors: Richard J. Young, Alice M. Matthews, Brach Poston
Format: Article
Language:English
Published: MDPI AG 2025-05-01
Series:Algorithms
Subjects:
Online Access:https://www.mdpi.com/1999-4893/18/5/296
Tags: Add Tag
No Tags, Be the first to tag this record!

Similar Items