Perspectives on the Use of Electronic Health Record (EHR) Data for Health Policymaking

EHRs have transformed healthcare by creating rich, longitudinal data sources that extend far beyond individual patient care. When analyzed appropriately, EHR data can provide timely insights into disease patterns, healthcare utilization, treatment effectiveness, and population health trends. These findings can inform evidence-based policymaking by helping decision-makers to evaluate existing policies, identify emerging public health needs, and better allocate resources. However, challenges related to data quality, privacy, interoperability, and representativeness must also be addressed. To further discuss this topic from a quantitative standpoint, we are happy to interview Dr. Sarah Lotspeich, Assistant Professor of Statistics at Wake Forest University.
Robert: Could you start off by telling us a little bit about yourself and your day-to-day work at Wake Forest? What do you enjoy most about your job?
Sarah: My job has a really nice balance, as my time is split pretty evenly between teaching and research (both of which I enjoy). Depending on the day of the week, my day-to-day can consist of classroom lectures, research meetings, and writing/coding time to myself. My 1:1 meetings with students to discuss their research projects are some of my favorite parts of the job, as I get to see them grow as (bio)statisticians and get excited about the field and the possibilities for them within it (especially since I work primarily with undergraduate students). My colleagues and I really like to have lunch and grab coffee together, too! So that’s a nice time in the day to chat about work and life.
Robert: What opportunities does Wake Forest offer for students or others who are interested in health policy?
Sarah: One of the things that impresses me most about the curriculum at Wake Forest is how, since we are a liberal arts university, students are encouraged to take courses broadly across disciplines. This opens up possibilities for them to find health policy early, for example, through courses toward the Health Policy & Administration minor. I would have loved to start learning about epidemiology and other key topics as an undergraduate! There are also ample opportunities for students to get involved in health policy research on the traditional Reynolda Campus (e.g., learning about the data-driven elements with us in the Department of Statistical Science) and with the School of Medicine (e.g., focusing more on the implications of our analyses with faculty in the Department of Social Sciences and Health Policy).
Robert: Can you give me an example of an EHR-based study that you have collaborated on? What implications could the findings have on healthcare policy?
Sarah: A large part of my research for the past few years has been toward operationalizing a whole person health score called the Allostatic Load Index (ALI) in EHR data. The ALI is essentially a snapshot of cumulative “wear and tear” on a person’s body, and it’s been fairly well studied in controlled settings like prospective studies, where it’s been found to predict key health outcomes (e.g., mortality, mental health, cardiometabolic diseases) and to quantify disparities in overall health (e.g., between racial groups). However, the ALI is calculated from ten component biomarkers, many of which are informatively missing in EHR data, since they require non-standard labs. We have been working to find a scalable way to overcome missing data and embed a score like this in EHRs, with the end goal being monitoring patient well-being across entire healthcare systems. Having an EHR-embedded ALI would open up opportunities to better understand individual-, neighborhood-, and hospital-level health and inform patient care, resource allocation, and policy. For example, we could identify where in a hospital’s service area to recruit patients or implement services to target those most in need (e.g., with the worst ALI), or evaluate the efficacy of a new policy on improving overall health of the patient population and reducing disparities (e.g., by improving the ALI).
Robert: What are the key strengths and limitations of using EHR data to inform public health or healthcare policy decisions?
Sarah: I think that the widespread availability and convenience are definite strengths, as we can curate so many different EHR datasets, depending on what we are interested in investigating. However, I think that the underlying mechanics of how EHR data are collected add challenges to responsibly conducting our analyses and translating findings into action. Those challenges may not always be limitations, but they sometimes require novel methods development to answer the questions we are interested in.
Robert: If an EHR dataset underrepresents certain populations or healthcare settings, how can potential biases be addressed to ensure that subsequent policy recommendations are more generalizable?
Sarah: With EHR data, underrepresentation is not an if but a whom. We have to keep in mind that, while EHR data are routinely collected as part of clinical practice, patients have to overcome key barriers to engage in a particular healthcare system at all. This selection bias is one of many key challenges with responsibly using EHR data for research, and fortunately there are some neat statistical tools we can use to try to overcome it. For example, we can compare representation in EHR data to administrative data (like the Census) and use these differences to decide how to upweight underrepresented groups and downweight overrepresented groups to make the EHR data more representative. In policy research, failing to account for selection bias in EHR data can lead to imbalanced recommendations and inaccurate evaluations (i.e., ones that really only apply to the groups who were well-represented in your healthcare system), leaving underrepresented people underserved by downstream decision-making.
Robert: What safeguards would you recommend to prevent unintended consequences of policy decisions that are based on EHR analyses?
Sarah: With EHR data, we have to remember that with “great-ish” data come great responsibility. There is a lot of exciting potential with EHR data, since they are increasingly available for us to re-purpose in research and policy. However, there is much-needed emphasis on the “re” in re-purpose, as EHR data are primarily collected to ensure patients are treated and billed appropriately. The clinicians inputting these data and programmers extracting them do not know what you may analyze from them at some point in the future, and there are not the same kinds of data monitoring safeguards in place as with controlled trials or prospective studies. Also, be mindful of the innate biases in HER data, including who we sample from (selection bias), what we observe about whom (informative missingness), and the possibility that health is imperfectly represented (measurement error).




Comments