For clinical operations leaders seamless data transfer today means pulling structured fields, such as vitals, labs, medications, demographics, straight from the EHR into the EDC, no re-keying required. It sounds like the finish line. Fewer transcription errors, faster query resolution, cleaner data.
It isn’t the finish line. It’s the on-ramp.
The Data You’re Not Capturing Is the Data That Matters Most
Here’s the uncomfortable math: by most industry estimates, somewhere between 45-70% of trial-relevant variables exist in semi- and un-structured formats – physician progress notes, discharge summaries, pathology reports, radiology impressions, nursing narratives. The fields that map cleanly into EDC forms, the ones most tools are built to move, represent a fraction of what’s actually in the chart.
That means a “complete” EHR-to-EDC transfer built only on structured fields is, by definition, incomplete. An adverse event buried in a physician’s note. A concomitant medication mentioned in a discharge summary but never entered as a discrete med order. A radiologist’s impression that contradicts the structured diagnosis code sitting one tab over.
If your data transfer solution can’t see that information, your trial can’t either.
Semi-Structured Data Is Where the Gap Quietly Widens
It’s tempting to think of this as a binary: structured fields on one side, free-text notes on the other. In practice, a huge volume of clinical data lives in between. Semi-structured content like lab reports with embedded reference ranges, scanned forms, flowsheets, and templated notes that follow a pattern but aren’t discrete fields in the database.
Semi-structured data is easy to overlook precisely because it looks almost structured. It has headers, sections, recognizable formats. But most source to sponsor pipelines either flatten it into unusable blobs of text or drop it entirely, because it doesn’t fit the rigid field-to-field mapping the tool was designed for. The result is a false sense of coverage: dashboards that report high match rates while quietly excluding an entire category of clinical evidence.
Why This Is Urgent Now, Not Eventually
This gap has always existed. What’s changed is how much it now costs to ignore it.
According to Tufts Center for the Study of Drug Development, the volume of data in clinical trials has grown by more than 6X in the last decade. Trials are more complex, with more endpoints, more real-world evidence requirements, and more regulatory scrutiny on data provenance. Sponsors and CROs are under pressure to move faster while regulators are asking harder questions about how source data was captured, transformed, and verified. FDA and international guidance increasingly emphasize traceability and human oversight, which is difficult to demonstrate when a meaningful share of source data was never systematically captured in the first place.
At the same time, the tools capable of extracting signal from unstructured and semi-structured clinical text — natural language processing, machine learning models trained on clinical narrative — have matured to the point where “it’s too hard to capture” is no longer a credible excuse. The technology gap has closed. The expectation gap has not.
Every month a trial runs on a structured-only data strategy is a month of accumulating exposure: missed safety signals, inconsistent source documentation, monitors doing by hand what software should have already surfaced, and a growing body of clinical evidence that technically exists at source, but functionally doesn’t exist for the trial.
A Solution Must Bring Holistic Data Intelligence
A data transfer strategy that only moves structured fields starts to tackle operational challenges, but a holistic solution is what truly moves the needle. Source to sponsor data transfer has to account for all three data types clinical documentation actually contains: structured, semi-structured, and unstructured.
The organizations that recognize this now and start asking their data transfer vendors pointed questions about how (or whether) they handle all data structures will be the ones building trials for scalability and maximizing results. The organizations that don’t will find out the hard way how quickly partial data transfer can set them back.