Ian Klosowicz

The hardest part of starting a portfolio project usually isn't the SQL or the dashboard. It's finding a dataset that's worth building on — one that has a real question in it, that doesn't look like every other portfolio you've seen, and that you can actually talk about in an interview without sounding like you followed a tutorial.
This post covers where to find good datasets, what makes a dataset portfolio-worthy, and how to tell quickly whether a dataset is going to produce something useful or stall out 3 hours in.
Not every publicly available dataset is worth building a portfolio project on. The ones that work share a few properties.
It has a real question inside it. Some datasets are inherently analytical — they contain enough structure and variance that you can ask something specific and get a specific answer. Others are descriptive registers that just list things. A dataset of every US city's population is a register. A dataset of every US city's population plus median income, housing costs, and job growth over 10 years has multiple real questions inside it.
It has at least 2 joinable tables, or can be made into them. Single flat files produce flat analysis. A dataset that's already normalized into multiple related tables, or one you can split into a fact table and at least 1 dimension table, lets you demonstrate data modeling skills. That signal matters to hiring managers reviewing both SQL and BI projects.
It isn't the dataset everyone uses. Titanic, Superstore, Airbnb NYC, AdventureWorks, and the Iris dataset all fail this test. The moment a reviewer recognizes the dataset, the project is mentally categorized as coursework. A dataset the reviewer hasn't seen before buys you the benefit of the doubt on analytical intent.
You can explain why the question matters. The best interview moments in a portfolio walkthrough are when a candidate says something like "I picked this dataset because I used to work in retail logistics and I wanted to understand the relationship between fulfillment speed and repeat purchase rate." That context signals genuine engagement with the problem, not just a search for data to fill a template.
Government open data is underused by candidates and consistently produces strong portfolio projects. The data is real, the questions are consequential, and the datasets are almost never tutorial datasets.
data.gov is the US federal government's primary open data portal. It has hundreds of thousands of datasets across agencies: transportation, health, education, housing, environment, finance. The quality varies, but the breadth means there's something for almost any analytical interest. Search by topic, not by dataset name.
City and county open data portals. Most major US cities publish open data: building permits, 311 service requests, crime incidents, business licenses, parking violations, transit ridership. These are particularly useful because they're geographically grounded — you can build a project about a city you know or are targeting in your job search, which makes the interview conversation more natural. Chicago, New York, San Francisco, Los Angeles, Seattle, and most other major metros have active portals.
census.gov and data.census.gov. The US Census Bureau publishes population, income, housing, employment, and demographic data at the national, state, county, and census tract level. Joining Census data to other local datasets — city permit data, health outcomes, school performance — produces the kind of multi-source analysis that reads as real research rather than a tutorial exercise.
healthdata.gov and CMS data. The Centers for Medicare and Medicaid Services publishes hospital performance data, provider utilization, and healthcare cost data at the facility and county level. If you're targeting healthcare analytics roles, this data is the right place to build. It has real questions inside it about quality, cost, and access that are actively studied in the field.
Bureau of Labor Statistics (bls.gov). Employment, wages, inflation, and industry data. The Consumer Expenditure Survey is particularly underused — it tracks spending by income bracket and category across thousands of households and is a natural fit for fintech, retail, or consumer analytics projects.
The best portfolio datasets are often industry-specific ones that align with where you're trying to work. If you're targeting a sector, use data from that sector.
Real estate: Zillow Research publishes housing price indices, inventory data, and rental data at the metro and zip code level. Redfin publishes similar data with different granularity. County assessor offices in most states publish property transaction records. These datasets support natural questions about market trends, affordability, and inventory dynamics.
Retail and e-commerce: The USDA publishes food retail data. The Census Bureau's retail trade surveys cover sales by store type over time. Some companies publish anonymized transaction datasets through research partnerships — the UC Irvine ML Repository has several retail transaction datasets that are real but not tutorial staples.
Healthcare: In addition to CMS data, the CDC publishes BRFSS survey data (behavioral risk factors at the state level), WONDER mortality data, and county health rankings. State health departments often publish facility-level data as well. These datasets support questions about outcomes, disparities, and resource access.
Finance: The SEC EDGAR system publishes financial filings from all publicly traded US companies. FRED (Federal Reserve Economic Data) publishes macroeconomic time series. These are used by actual financial analysts and produce projects that read as real work in that context.
Education: The National Center for Education Statistics publishes school-level data on enrollment, test scores, demographics, and funding. State departments of education often publish district-level data as well. If you're interested in public policy or edtech analytics, these are the right sources.
Sports data is some of the most accessible, cleanest, and most naturally analytical data available. It also makes for interviews that are easy to talk about because the questions are intuitive and the findings are interpretable without domain expertise.
Baseball Reference (baseball-reference.com) has play-by-play, season-level, and career-level data going back over a century. The Lahman Database is a downloadable version of much of this data in relational table format — multiple tables, foreign keys, joinable across seasons and players. This is one of the few publicly available datasets that comes already normalized.
Basketball Reference (basketball-reference.com) covers NBA and WNBA player and team statistics. FBref covers global soccer. Pro Football Reference covers NFL. All of these have tabular data downloadable at the game, season, and career level.
Stathead and Sports Reference APIs allow more structured access for people who want to build SQL queries against the data directly rather than downloading flat files.
Spotify and music data. Spotify's public API provides audio features for tracks (danceability, tempo, energy, valence) and is frequently used for portfolio projects. The advantage: it's genuinely interesting data. The disadvantage: it's become common enough in portfolios that some reviewers have seen it frequently. Use it if you have a genuinely specific question, not just to visualize audio features.
Financial data is high-signal for roles in finance, fintech, and business analytics generally. It also tends to be well-structured and naturally multi-table.
SEC EDGAR. Every public company files financial statements with the SEC. EDGAR makes these available in structured format. You can pull income statements, balance sheets, and cash flow statements for any publicly traded company over multiple years and build analysis around revenue trends, profitability, or sector comparisons.
Crunchbase public data. Crunchbase publishes a subset of startup funding data publicly. Questions about funding rounds by sector, company stage, and geography produce interesting analysis that reads as relevant to venture capital, startup operations, or investment analytics roles.
World Bank Open Data. Country-level economic, health, and education indicators going back decades. Good for macro-level analysis comparing development trajectories, regional trends, or the relationship between economic indicators and outcomes like life expectancy or education attainment.
FRED (Federal Reserve Economic Data). Time series data on interest rates, inflation, unemployment, GDP, housing starts, and hundreds of other macroeconomic indicators. Naturally suited to time-series analysis projects and reads as directly relevant to finance and economics roles.
Kaggle has a reputation for tutorial datasets, but the platform has tens of thousands of datasets, most of which have never appeared in a tutorial. The problem is that the most popular ones — Titanic, House Prices, the Iris classification set — are the first results in a search and the ones that show up in every beginner course.
Using Kaggle well means going past the first page of results. Search by topic rather than by "beginner" or "popular." Look for datasets with real provenance — ones that were collected from a real source, not generated for a competition. Sort by "Most Votes" in a specific category to find datasets the data science community finds valuable, then check whether those datasets appear in tutorials before building on them.
Kaggle also hosts datasets from companies that have published data for transparency or research purposes: Airbnb (a different slice than the standard NYC tutorial dataset), Uber movement data, and various e-commerce and logistics companies. These are real datasets with real analytical questions inside them and don't carry the tutorial-dataset stigma.
Spending 3 hours on a dataset only to discover it doesn't have enough structure for the project you had in mind is a common and avoidable problem. Run this quick evaluation before committing:
If the dataset passes all 5, it's worth building on. If it fails 2 or more, look for something else rather than trying to force it.
The single strongest differentiator in a portfolio interview is domain knowledge that most other candidates don't have. An analyst who worked in hospitality before transitioning to data and builds a project on hotel revenue management data will talk about that project differently than someone who picked it because it was available. The context is visible, the question is more specific, and the findings are more interpretable.
If you have industry background — healthcare, retail, finance, logistics, education, nonprofits, construction, manufacturing — look for public data from that industry first. You already know what questions matter in that context. You already know what a good answer looks like. That contextual knowledge will come through in every walkthrough, and it's something that candidates who just picked a dataset because it was there can't replicate.
I came from a background outside data and built projects on data that I found genuinely interesting — not because the datasets were the most technically impressive, but because I could speak to why the questions mattered. That genuine engagement carried further in interviews than technical polish would have on its own.
If you want a structured approach to picking datasets, scoping the right questions, and building projects that read as real analytical work, the Analyst Hive program covers all of it in Month 1 as part of the 3-project build sequence.
How do I find datasets that haven't been used in tutorials?
Search by topic on government portals and industry-specific sources rather than searching for "datasets for data analysts." The tutorial datasets show up when you search for learning resources. Real-world data shows up when you search for the topic itself: city permit data, hospital performance, sports statistics by season, retail trade surveys. Cross-check a dataset you're considering by searching it in Google alongside words like "tutorial" or "portfolio project" — if it shows up in multiple tutorials, find something else.
Is it okay to use a Kaggle dataset for a portfolio project?
Yes, if it's not a standard tutorial dataset and you have a real question for it. The Titanic and House Prices datasets should be avoided. Thousands of other Kaggle datasets have never appeared in a tutorial and are perfectly usable. Filter by topic and look for datasets that come from real sources rather than ones created specifically for competitions or practice.
What if the dataset I find is messy or incomplete?
That's usually a feature, not a problem. Messy, real-world data gives you the opportunity to show data cleaning skills — handling nulls, standardizing formats, removing duplicates, dealing with inconsistent categories. Document what you found and what you did to address it in your README or dashboard methodology section. Candidates who demonstrate data prep skills alongside analysis skills are stronger than candidates who only built on clean data handed to them.
How many rows does a dataset need to be useful?
10,000 rows is a reasonable minimum for most analytical questions. Below that, segment-level aggregations become unreliable and the findings are harder to trust. For time-series analysis, you want at least 2 to 3 years of data at the granularity you're working at. For geographic analysis, you want enough coverage that regional comparisons are meaningful. If the dataset you want is smaller, consider joining it to a second source to increase the analytical surface area.
Should I use real-time or live data in my portfolio project?
Not unless you have a strong technical reason to. Live data introduces reliability problems: the API can change, the data format can shift, and the dashboard can break in the middle of an application review. For portfolio purposes, a static snapshot of good data is more reliable than a live feed of data that might stop working. If you want to demonstrate API skills, a recorded walkthrough of a working live connection is more portfolio-safe than a link that depends on an external service staying stable.
Can I use data from my current or previous job?
Only if it's fully anonymized or mocked and you're certain your employer has no objection. Using actual company or client data in a public portfolio is a risk that can have professional and legal consequences. The safer approach: rebuild the analysis structure on public data that approximates what you worked with. You can describe the real-world context in interviews without exposing proprietary data publicly.
If you want a structured approach to finding the right dataset, scoping the right question, and building a project that reads as real analytical work rather than a tutorial, the Analyst Hive program walks through all of it in Month 1. Daily tasks, built around getting hired.