Organizations have access to more data than ever before. The challenge has shifted from finding data to making sense of it.

Understanding where programs are already making an impact is one part of the picture. The next step is looking ahead to where additional outreach and resources could make a difference. For our analytics initiative supporting the Kay Yow Cancer Fund, we brought together public health data from various sources to identify locations for community outreach, education and cancer screening programs.

The data came in many formats from different federal health agencies: Excel files, flat files, ZIP archives, downloadable CSVs and REST APIs. Each source had its own structure, update cycle, and technical requirements. Luckily, all sources published their data at the county level, so we had a common key to bring them together.

Individually, these datasets provide valuable insights. Together, they can reveal something more useful: where cancer burden, social vulnerability, screening access and existing resources intersect. Simply downloading and integrating these diverse data types can quickly overwhelm organizations, especially those without a dedicated data engineer or someone experienced with data.

For this project, we used SAS® Viya® to integrate multiple sources for analysis and visualization.

Read: Behind every data point is a person: Turning insight into impact with the Kay Yow Cancer Fund

The public health data challenge

Many health datasets are publicly available, but their formats vary. We incorporated several sources for this project. Each source provides a different view into overall community health.

ACS data helps quantify population characteristics and demographics. CDC SDOH data highlights factors such as economic conditions, education, transportation and health care access. NCI data show where cancer incidence and mortality rates are high and increasing or falling. The FDA mammography facility provides physical screening locations to help address breast cancer prevention needs and identify service gaps. We also leveraged data on Kay Yow grant recipients to show where community investments already exist.

No single dataset can tell the Fund where outreach is needed most. Looking at them together, however, begins to show where high cancer burden may overlap with social vulnerability, lower screening rates and gaps in access to care.

For the Kay Yow Cancer Fund, we set out to create an analytical and interactive dashboard to identify counties with high rates of different types of cancer, high levels of social vulnerability and low rates of cancer screening and prevention to plan EmPOWERment Tours. The Tours provide communities with cancer education, screening awareness and access to local health resources.

The dashboard provides the Fund with another way to identify communities where outreach, education and screening resources could have the greatest impact.

Connecting data across formats

Answering that question first required us to bring very different types of data into the same environment.

For this project, we used several techniques to ingest the data:

  • We downloaded and processed Census API data using PROC HTTP.
  • We retrieved CDC datasets directly from downloadable CSV endpoints.
  • We read ZIP archives published by the FDA and other Federal Agencies.
  • We loaded all datasets into SAS Viya for scalable analytics and visualization.

The code and scripts are relatively straightforward because SAS provides the capabilities needed to interact with modern data sources, such as REST APIs and less modern but still valuable ones, such as flat files.

What I love the most about SAS is that, as a data analyst, I can develop a script that reads in these data sources directly into the platform for integration, analysis and visualization. I don’t have to download a file, upload it to a server and then import it to SAS. All I need to do is save my script program file and I can rerun the code if I change environments or the data refreshes.

As a bonus, I can keep all my scripts in a single GitHub repository that integrates with SAS Data and AI Studio – my personal preference for a SAS programming IDE – and with VS Code.

Back to the coding techniques I used in this project.

PROC HTTP lets me interact directly with web APIs and data services without requiring additional tools. It supports standard HTTP methods, including GET and POST, and lets me access REST APIs and web services directly from SAS programs.

After more than 15 years of programming in SAS, I think PROC HTTP may be my favorite PROC.

The following shows how to access the Census ACS data in JSON format as an example. SAS has a purpose-built access engine that reads JSON files and presents them like a traditional SAS table.

The code uses the following SAS procedures and statements:

  • FILENAME statement
  • PROC HTTP
  • LIBNAME statement with the JSON engine
  • DATA step

ACS data is very wide, with non-descriptive variable names, so after this initial ingest step, we need to perform post-processing and apply variable labels. Luckily, the JSON library also includes descriptive variable labels for each variable code, making it a convenient one-stop shop!

If you want to run this code, you will need to request your own API key

Similarly, federal agencies sometimes publish compressed files. SAS provides FILENAME ZIP functionality that allows programs to inspect and extract ZIP contents programmatically within SAS workflows. This blog post by Chris Hemedinger provides an excellent example of using FILENAME ZIP to access data files stored in ZIP archives without relying on external utilities. I used his approach to develop the script below for reading FDA mammography facility data.

The following code demonstrates how to download a compressed file from the FDA and use the SAS ZIP filename method to read mammography facility data directly from the archive.

The code uses the following SAS procedures and statements:

  • Macro variables and functions
  • FILENAME statement
  • PROC HTTP
  • DATA step
  • INFILE and INPUT statements
  • INFORMAT, FORMAT and LABEL statements

These capabilities allow data engineers and analysts like me to spend less time wrestling with file formats and more time understanding what the data can tell us.

Creating a common geographic language

Bringing the datasets into one environment solved one problem. Making them speak the same geographic language created another.

Every data source can describe geography differently. One data set may contain county names. Another may contain five-digit ZIP Codes. A third may provide latitude and longitude coordinates. Others may use state and county FIPS codes.

Without standardization, we can’t accurately combine the datasets.

In this project, we chose the county FIPS code as our common key. We then used text functions to concatenate the state and county FIPS codes to create our key codes and a unique geographic identifier.

That common identifier helped us to analyze demographic data, social determinants, metrics, health care access information and Kay Yow Cancer Fund grant recipient information together at the county level.

Now we could move beyond isolated statistics and see how different conditions overlap within the same community.

Bringing it all together with mapping and visualization

After we connected the datasets, we could look at demographic indicators, social determinants, health care resources and the Fund’s existing investments together.

That’s where the data starts becoming useful for the Fund.

Rather than reviewing cancer rates, screening access, social vulnerability and existing investments separately, the Fund can ask how those factors intersect within individual communities.

For example:

  • Which counties have elevated social risk indicators, high cancer incidence rates and low cancer screening rates?
  • Do those counties have local or mobile mammography facilities?
  • Which communities with low cancer screening rates sit near Kay Yow Cancer Fund grantees?
  • Where might additional outreach programs have the greatest impact?

Those questions shift the analysis from understanding individual datasets to understanding communities.

Figure 1: Using the dashboard, the Kay Yow Fund can now interact with county Cancer Incidence and Mortality rates alongside their Grantees, while also considering other community health factors such as the CDC Social Vulnerability Index, Mammogram rates, cigarette smoking rates, dentist and doctor visit rates, and economic factors such as the unemployment rate.

The bigger opportunity

While this example focuses on the Kay Yow Cancer Fund’s mission to improve cancer screening and community outreach, the underlying lesson applies much more broadly.

Important questions rarely live in one single data set.

Understanding where a community needs support may require organizations to consider health outcomes alongside demographics, economic conditions and access to existing services and resources. Any one of those factors provides only part of the picture.

For the Kay Yow Cancer Fund, bringing them together creates a stronger understanding of where programs already make an impact and where future community outreach efforts might provide the greatest benefit.

That’s also how this work builds on what came before it. Measuring impact helps us understand what programs have already accomplished. Connecting that evidence with the conditions communities face can help inform what happens next.

The goal is not simply to collect more data. The goal is to connect the right data so organizations can see what they couldn’t before and make better decisions as a result.

Support the Kay Yow Cancer Fund

Help advance the Kay Yow Cancer Fund’s mission to improve outcomes for women facing cancer through research, education and access to care.

Learn more about the Kay Yow Cancer Fund →

Hear more perspectives on health care

Join us for The Health Pulse podcast as we explore fresh perspectives on digital transformation, data and innovation across health care and life sciences.

References




Source link


administrator