
Listen instead
Three physicians spent hours verifying 2,710 individual data points extracted from 20 complex patient charts to test a basic clinical premise: can software reliably read and parse a doctor's hand-typed progress notes? The manual audit, detailed in a new pre-print validation manuscript, represents the high-stakes validation work required to move artificial intelligence past simple administration and into clinical research.
Historically, medical researchers looking to track how real patients respond to new therapies have faced a stark choice. They could either rely on structured billing codes, which lack clinical detail, or hire human abstractors to manually read thousands of pages of progress notes. This manual review process is the single greatest bottleneck in clinical research, often requiring clinical research organizations to employ teams of nurse-abstractors to painstakingly transcribe data into registries. Now, a multi-institutional research team has demonstrated an automated method that could unlock this unstructured text while preserving the strict audit trails that medical researchers and regulatory bodies require.
The validation study, co-authored by researchers from RespondHealth, Drexel University, Stanford University, and the University of Miami, introduces a clinical large-language-model framework. The system is designed to extract complex longitudinal data from unstructured progress notes while maintaining verifiable, sentence-level provenance. By linking every extracted data point directly to a specific source sentence in the patient record, the researchers aim to solve the trust problem that has long limited the clinical adoption of artificial intelligence in regulatory-grade real-world evidence.
Verifiable Precision in Clinical Extraction
To evaluate the accuracy of the automated system, the research team subjected the model's outputs to rigorous manual validation. Three independent physicians audited 2,710 data extractions across 20 sampled patient charts. The results showed high accuracy levels. The framework achieved 99.5 percent precision for extracting clinical phenotypes and 95.5 percent precision for extracting medication details.
However, high precision alone is insufficient for clinical research. In medical environments, regulators and researchers must be able to verify the underlying source of any automated finding. Traditional chart abstraction methods are prone to human fatigue and transcription errors, meaning that even manual registries require secondary audits. The framework addresses this requirement by retaining traceable source-note evidence for every single extraction.
By providing sentence-level provenance, the model allows human auditors to click on an extracted clinical data point and immediately view the exact sentence in the progress note from which it was derived. This capability transforms the role of clinical abstractors. Instead of spending hours hunting through clinical narratives to find symptoms, dosages, and side effects, human specialists can act as high-level editors, verifying automated extractions in a fraction of the time. This shifts the bottleneck from manual search to rapid verification.
Tracking Real-World GLP-1 Performance

To demonstrate the practical value of the technology, the researchers applied the framework to a large clinical dataset. They analyzed the unstructured progress notes of 16,061 adults who were initiating treatment with injectable GLP-1 receptor agonists, specifically semaglutide or tirzepatide. By extracting clinical data directly from doctor notes, the model mapped detailed real-world weight-loss trajectories that are typically invisible in structured billing codes.
The resulting analysis revealed that real-world weight-loss trajectories differed based on a patient's baseline glycemic status. For patients with a normal baseline HbA1c, the model estimated a 12-month weight loss of 7.7 percent, with a median time of 210 days to reach a 5 percent weight loss. In contrast, among patients with poorly controlled diabetes, the estimated 12-month weight loss was only 2.7 percent, and the median time to reach a 5 percent weight loss stretched to 413 days.
The model also identified demographic differences in weight-loss outcomes. Among female patients, the predicted 12-month weight change was 6.1 percent, compared to 4.0 percent among male patients. Age also influenced the outcomes. Adults aged 20 to 39 had an estimated 12-month weight change of 8.1 percent, while adults aged 40 and older had an estimated change of 5.1 percent.
While these findings are based on an un-peer-reviewed pre-print, they illustrate the type of granular, longitudinal insight that remains locked within free-text medical records. Clinical trials show drug performance under idealized conditions, but real-world tracking shows how these therapies perform in diverse, non-adherent patient populations.
Transforming Post-Market Surveillance

For health systems and biopharma sponsors, the study highlights an opportunity to modernize real-world evidence collection. Currently, most retrospective studies rely on ICD-10 diagnostic codes. These codes are designed for billing rather than clinical precision. They frequently miss critical elements of the patient journey, such as specific side effects, precise dose titrations, and early indicators of drug tolerability. These nuances are typically documented only within the narrative text of progress notes.
By transforming unstructured narratives into traceable, computable data, this artificial intelligence framework allows organizations to conduct post-market surveillance and retrospective analyses at a scale that was previously impossible. Biopharma companies can use these tools to defend formulary placement with payers by demonstrating real-world value, while health systems can identify patient cohorts eligible for specialized clinical trials.
The ultimate goal is not to eliminate human oversight, but to reduce the burden of manual chart abstraction. While formal peer review for this manuscript is still pending, the initial validation results offer a practical blueprint for modern clinical data management. By solving the provenance issue, the researchers have addressed the main obstacle to scaling artificial intelligence in clinical research, ensuring that automated insights remain fully auditable.
Source: The HealthTech Signal

