Ian Klosowicz

Healthcare is one of the strongest industries to target if you're building a data analyst portfolio. The public data is real, the questions are consequential, and the analytical problems map directly to what healthcare analytics teams work on every day. A well-built healthcare project signals domain interest and stands out from the sea of retail sales dashboards in ways that generic datasets can't.
Understanding what healthcare data analysts actually work on day to day helps you frame these projects in the direction hiring managers in that space will recognize.
This post covers 8 specific project ideas using publicly available healthcare data, what analytical question each one answers, what skills each one demonstrates, and where to find the data to build them.
Before picking a project, it helps to know what the work actually looks like in healthcare. Most healthcare analyst roles sit inside hospital systems, health insurance companies, pharmacy benefit managers, digital health companies, or public health agencies. The analytical work falls into a few consistent buckets:
A portfolio project that maps to any of these areas signals to a hiring manager that you understand the context of the work, not just the tools. That contextual awareness is something most entry-level candidates don't demonstrate.
Most of the data you need for these projects is publicly available for free. The major sources:
CMS (Centers for Medicare and Medicaid Services) publishes hospital-level data on readmission rates, complications, patient experience, payment penalties, and provider utilization. The Hospital Compare dataset is the most commonly used. Access at data.cms.gov.
CDC WONDER publishes mortality data by cause, age, geography, and year. It's searchable and downloadable in structured format. Access at wonder.cdc.gov.
BRFSS (Behavioral Risk Factor Surveillance System) is an annual CDC survey covering chronic disease prevalence, health behaviors, and risk factors at the state level. Access at cdc.gov/brfss.
County Health Rankings is published annually by the University of Wisconsin Population Health Institute in partnership with the Robert Wood Johnson Foundation. It ranks every US county on health outcomes and health factors. Access at countyhealthrankings.org.
AHRQ (Agency for Healthcare Research and Quality) publishes the Healthcare Cost and Utilization Project (HCUP), which contains state-level inpatient and emergency department data. Some datasets require registration but are free. Access at ahrq.gov/data.
healthdata.gov aggregates federal health datasets from multiple agencies and is a useful starting point for discovery.
1. Hospital readmission rates and the factors that predict them
Question: Which hospital characteristics are most strongly associated with 30-day readmission rates, and do the same factors predict readmissions across different conditions?
Data: CMS Hospital Compare readmission data, joined to hospital-level characteristics (bed count, teaching status, ownership type, urban/rural classification).
What it demonstrates: multi-table SQL joins, aggregations at the facility level, trend comparison across condition types, a Power BI or Tableau dashboard with filter by condition and facility type. The data model has a clean fact/dimension structure: readmission rates as the fact, hospital characteristics as the dimension.
Why it reads as real: 30-day readmission rates are a CMS quality measure tied to payment penalties. Hospitals actively track and try to reduce them. This is the kind of analysis a healthcare system analytics team runs regularly.
2. County-level health outcome disparities
Question: Which counties have the largest gap between health outcomes and their socioeconomic conditions, and what does that gap suggest about where interventions would have the highest return?
Data: County Health Rankings, joined to Census income and poverty data by FIPS code.
What it demonstrates: multi-source data join on a geographic key (FIPS code), gap analysis, ranking with window functions, a choropleth-style dashboard or ranked comparison view. Good for showing that you can work with geographic joins and multi-source analysis.
Why it reads as real: population health analysts at payers and health systems run exactly this analysis to identify counties or regions where targeted outreach would reduce costs and improve outcomes.
3. Preventable hospitalization trends by state
Question: How have preventable hospitalization rates changed across states over the last decade, and which states have improved most relative to their starting point?
Data: AHRQ Prevention Quality Indicators (PQIs) by state and year. These are specifically designed to measure hospitalizations that could have been avoided with better ambulatory care.
What it demonstrates: time series analysis, period-over-period change calculations, ranking states by improvement rate rather than absolute level, a line chart trend dashboard. Window functions for year-over-year change are natural here.
Why it reads as real: PQIs are a standard metric in public health and health system strategy. A candidate who knows what a Prevention Quality Indicator is signals genuine familiarity with the field.
4. Chronic disease prevalence and primary care access
Question: Is there a relationship between primary care physician density and chronic disease prevalence at the county level, and which counties have the worst combination of high disease burden and low access?
Data: County Health Rankings (provider ratios, disease prevalence) joined to BRFSS chronic condition data.
What it demonstrates: correlation analysis at the county level, scatter plot or bubble chart visualization, segment analysis identifying high-burden / low-access counties, SQL aggregations and joins across 2 public sources.
Why it reads as real: this is a classic population health and health equity analysis. Payers, health systems, and public health agencies all run versions of this to inform where to open clinics, deploy telehealth, or target chronic disease management programs.
5. Medicare spending variation across hospital referral regions
Question: How much does Medicare spending per beneficiary vary across hospital referral regions, and is higher spending associated with better outcomes or just higher utilization?
Data: Dartmouth Atlas of Health Care publishes Medicare spending and utilization data by hospital referral region. CMS also publishes regional Medicare data.
What it demonstrates: geographic aggregation, spending vs. outcome correlation, ranking regions by efficiency (outcome per dollar), a dashboard with regional comparison and drill-down. This is a slightly more sophisticated analytical question that shows you understand the difference between input (spending) and output (outcomes).
Why it reads as real: Medicare spending variation is one of the most studied questions in health policy. The Dartmouth research showing that higher spending doesn't correlate with better outcomes is foundational to value-based care discussions. Knowing this context will come through in the interview.
6. Drug overdose mortality trends by demographic and geography
Question: How have drug overdose mortality rates changed across states and demographic groups since 2015, and which combinations of geography and demographics show the steepest trajectories?
Data: CDC WONDER mortality data filtered to drug overdose causes of death (ICD-10 codes X40-X44, X60-X64, X85, Y10-Y14). Downloadable by state, year, age group, and race/ethnicity.
What it demonstrates: filtering and aggregating CDC mortality data, time series by multiple dimensions, identifying inflection points (fentanyl transition around 2016-2017), a multi-series line chart dashboard with demographic filters. SQL filtering on ICD-10 code ranges is a real skill in healthcare analytics.
Why it reads as real: public health agencies, payers, and health systems actively track overdose trends for resource allocation and program design. This analysis appears in real public health dashboards at the state level.
7. Hospital patient experience scores and their operational drivers
Question: Which HCAHPS patient experience dimensions (communication with nurses, responsiveness, cleanliness) are most strongly correlated with overall hospital ratings, and does that relationship hold across hospital size and ownership type?
Data: CMS HCAHPS patient experience survey data, joined to hospital characteristics from Hospital Compare.
What it demonstrates: correlation analysis across multiple survey dimensions, segmentation by hospital type, a dashboard that lets a viewer explore which experience factors matter most, SQL joins across CMS datasets. HCAHPS is the standard patient experience measure across US hospitals — knowing the acronym signals domain familiarity.
Why it reads as real: patient experience is tied to CMS reimbursement through the Value-Based Purchasing program. Hospital operations teams actively work to improve HCAHPS scores. This is operational analytics work.
8. Health insurance coverage gaps before and after ACA expansion
Question: How did uninsured rates change across states that expanded Medicaid under the ACA versus those that didn't, and did expansion states see corresponding changes in preventable hospitalization rates?
Data: Census Bureau American Community Survey (insurance coverage by state and year), AHRQ Prevention Quality Indicators, state Medicaid expansion dates (publicly available list).
What it demonstrates: policy impact analysis using a quasi-experimental design (expansion vs. non-expansion states as a natural comparison group), multi-source join, time series before and after a policy change date, a dashboard that shows divergence between the 2 groups over time. This is a more analytically sophisticated project that goes beyond descriptive analysis into causal reasoning.
Why it reads as real: health policy analysts, payers, and public health researchers have published extensively on ACA expansion effects. Building this analysis from primary data shows you understand the policy context and can work with real longitudinal data.
If you want a structured framework for picking the right project from this list, scoping the question, and building it to a standard that clears the hiring bar, the Analyst Hive program walks through exactly that in Month 1.
The domain context is half the value of a healthcare project. Use it.
When you walk through a healthcare portfolio project in an interview, explain what the metric actually is and why it matters to a healthcare organization, not just what the data shows. "30-day readmission rates are a CMS quality measure that's tied to payment penalties under the Hospital Readmissions Reduction Program — so reducing them isn't just a quality goal, it's a financial one" is a sentence that signals genuine understanding of the field. Most entry-level candidates can't say that.
If you have a background in healthcare — clinical, administrative, insurance, pharmaceuticals, or adjacent — connect the project explicitly to work you've seen or problems you've encountered. That personal connection is the most powerful thing you can bring to a portfolio walkthrough, and it's something no one else in the candidate pool can replicate.
If you don't have a healthcare background, pick the project whose question you find most interesting and spend an hour reading about the real-world context before the interview. You don't need to be a domain expert — you need to be able to explain why the question matters to the organization that would be asking it.
Do I need healthcare experience to build these projects?
No. All of these use publicly available data that anyone can access. Healthcare background helps you choose the right question and speak to the context in interviews, but it's not required to build the project. If you don't have a healthcare background, pick a question that's easy to explain in business terms and spend time understanding the real-world stakes of the analysis before any interview.
Which of these projects is best for a Power BI portfolio?
The hospital readmission project (#1) and the HCAHPS patient experience project (#7) both have clean data models with natural fact/dimension structures that translate well to Power BI. They also have enough categorical dimensions (condition type, hospital size, ownership type, region) to build a multi-page interactive dashboard with meaningful filters. Either of those works as a strong first BI portfolio piece in the healthcare space.
Which project is best for a SQL portfolio?
The county health disparities project (#2) and the preventable hospitalization trends project (#3) both require multi-source joins, date-based aggregations, and window functions for ranking or year-over-year change. Either translates well to a SQL-focused GitHub project with 3 to 5 queries, each answering a sub-question, with a README that explains the findings.
Is CMS data hard to work with?
It's real data, which means it has some messiness. Column names are verbose, some fields require lookup tables to interpret, and joining across CMS datasets requires matching on provider numbers (CCN) rather than names. None of that is hard, but it's not as clean as a tutorial dataset. Document how you handled the messiness in your README — that documentation is part of the portfolio signal.
What if I want to work in health insurance rather than a hospital system?
The Medicare spending variation project (#5), the ACA expansion project (#8), and the county-level chronic disease project (#4) all map most naturally to payer analytics work. Insurance companies and health plans care about spending variation, risk stratification, and the relationship between access and outcomes. Lead with those if health insurance is your target.
Can I combine 2 of these into one larger project?
Yes, and it often strengthens the result. Combining county health rankings with preventable hospitalization data, or HCAHPS scores with readmission rates, creates a richer dataset and a more interesting analytical question. Just make sure the combined question is still specific enough to answer clearly — a broader dataset doesn't help if the question becomes too vague to drive a specific finding.
If you want a structured path through building your first healthcare project and getting it onto a resume that actually gets read, the Analyst Hive program covers the build, the framing, and the walkthrough prep. Daily tasks, built around getting hired.