false
OasisLMS
Login
Catalog
OnDemand: AI for Obesity Research and Clinical Pra ...
Generating AI-Ready Datasets for Biomedical Resear ...
Generating AI-Ready Datasets for Biomedical Research
Back to course
[Please upgrade your browser to play this video content]
Video Transcription
Video Summary
Chris Hartshorn, an NIH expert in digital mobile technology and AI, discussed the real challenges of using artificial intelligence in biomedical research. He argued that current “AI” is often better understood as advanced data science, with most FDA-approved tools still limited to narrow, single-data-type uses such as radiology. Hartshorn emphasized that the biggest barrier is not the model itself, but preparing data so it is FAIR, well-documented, standardized, representative, and machine-readable. He warned that biomedical datasets are often noisy, incomplete, and biased, and that large language models add new risks such as hallucinations, confabulation, and even deliberate manipulation. <br /><br />He stressed that AI becomes most transformative when used on large, multimodal, multi-site datasets to generate new hypotheses, not just confirm existing ones. He also highlighted the need for better workforce training in “good algorithmic practice,” stronger standards for preprocessing, and clinical trials designed with AI in mind from the start. <br /><br />Hartshorn described NIH initiatives such as AIM AHEAD, Bridge to AI, and the Nutrition for Precision Health program as efforts to create AI-ready datasets and best practices. In Q&A, he urged caution with LLMs in clinical and research settings, recommending they remain closely monitored and not yet relied on for critical biomedical tasks.
Asset Subtitle
By Chris M. Hartshorn, PhD
Keywords
artificial intelligence
biomedical research
FAIR data
machine-readable datasets
large language models
NIH initiatives
data standardization
×
Please select your language
1
English