Ian Klosowicz

A tutorial project and a portfolio project are not the same thing. A tutorial project proves you followed instructions. A portfolio project proves you can think. Hiring managers have seen enough of both to tell the difference inside 30 seconds, and only one of them earns a callback.
The good news is the gap between them is smaller than it looks. You don't need a new dataset type or a more advanced technique. You need a different starting point: a real question instead of a dataset to explore. This post walks through what that shift looks like in practice, with concrete project ideas that read as analytical work rather than coursework.
Before fixing the problem, it helps to know exactly what signals tutorial work to a reviewer.
The dataset is universally recognized. Titanic survival data. Superstore sales. Airbnb NYC listings. AdventureWorks. These datasets exist specifically for learning, which means experienced reviewers have seen them hundreds of times. The moment one appears, the project is mentally categorized as practice.
There's no stated question. The project explores the data. It shows total sales, product breakdown, regional distribution. All accurate, none of it answering anything specific. Exploratory work is valuable internally. It doesn't function as a portfolio piece because it doesn't demonstrate analytical judgment.
The structure mirrors the tutorial. Dashboard with the same 4 charts from the course. SQL queries that demonstrate each function from the lesson, run in order. The work isn't wrong, but the shape of it is identifiable as instruction-following rather than problem-solving.
The findings describe the data instead of interpreting it. "Revenue is highest in the West region." "Most customers are between 25 and 44." Those are observations. A portfolio project leads somewhere: "West region revenue is highest but margin is lowest, driven by 2 product categories that account for 60% of regional discounting." That's an insight with a business implication.
None of these signals require advanced skills to fix. They require a different intent when you start.
The single change that turns a tutorial project into a portfolio project: write the question before you open the tool.
The question has to be specific enough that you can tell when you've answered it. "Analyze sales performance" is not a question. "Which product categories had the largest gap between Q3 and Q4 revenue, and what does the customer segment breakdown look like for those categories?" is a question. When you've answered it, the project is done. When you haven't, you know what's missing.
A strong portfolio question has 3 properties:
Start there. The tools, the charts, and the SQL follow from the question. The question is the project.
These project ideas are built around specific questions, not dataset types. Most use publicly available data.
Customer retention analysis on e-commerce transaction data. Question: What percentage of first-time buyers make a second purchase within 90 days, and does that rate vary by acquisition channel or product category? This requires cohort logic, date arithmetic, and window functions. It mirrors real retention analysis work and leads to a recommendation about where to focus retention efforts.
Layoff trend analysis using public tech layoff data. Question: Which industries and company stages have seen the highest layoff rates since 2022, and is there a relationship between company funding stage and layoff size? Publicly available datasets exist for this. The question is timely, the findings are specific, and the SQL requires joins, aggregations, and ranking functions.
City permit data analysis for a target metro. Question: Which neighborhoods have seen the largest increase in construction permits over the last 3 years, and what permit types are driving the growth? Most US cities publish permit data. This reads as real local government or real estate analytics work and uses date filtering, GROUP BY, and trend comparison.
Sports performance analysis on season-level data. Question: Which teams significantly outperformed their pre-season win projections, and does home/away split explain the variance? Sports reference sites have decades of data. The question is specific, the SQL involves joins across multiple tables (team stats, schedule, projections), and the findings are interpretable.
Job posting analysis using scraped or downloaded listings. Question: How do salary ranges and required skills differ between analyst roles at companies under 500 employees versus those over 5,000? This is directly relevant to the analytics job market and demonstrates that you understand the field you're entering. LinkedIn job data APIs and public datasets like Levels.fyi or H1B disclosure data make this buildable.
Each of these reads as work a practitioner would actually do, not a lesson someone assigned. The SQL requirements are real but achievable at the entry level.
The same principle applies to dashboard projects. Start with the question, not the dataset.
Public health outcome dashboard. Question: Which counties have the highest rates of preventable hospitalizations, and how does that correlate with access to primary care providers? CDC and CMS publish county-level data on both. This reads as healthcare analytics work, the data model requires joining 2 or more sources, and the findings have a clear policy implication.
Local real estate market tracker. Question: How have median sale prices and days-on-market changed across neighborhoods in a target city over the last 2 years, and which neighborhoods are accelerating versus cooling? Zillow, Redfin, and most county assessor offices publish this data. The dashboard structure maps to standard real estate analytics and the question is answerable from public sources.
College athletics financial analysis. Question: Which athletic programs have the largest gap between revenue and expenses, and does conference affiliation predict financial sustainability? The NCAA publishes financial data for all Division I programs. This is a non-obvious dataset that produces specific findings and reads as sports business analytics.
Retail foot traffic and revenue correlation. Question: Across store locations, does foot traffic index correlate with revenue per square foot, and which locations are underperforming relative to their traffic? Placer.ai publishes some foot traffic data; retail performance benchmarks are available through industry reports. This reads as retail analytics work.
Personal finance benchmarking tool. Question: How does average spending by category compare across income brackets, and which categories show the largest variance? Bureau of Labor Statistics Consumer Expenditure Survey publishes this data annually. This reads as fintech or personal finance analytics and is a dataset most reviewers haven't seen in a portfolio.
The common thread: each one uses a non-tutorial dataset, states a specific question, and leads to findings that imply a real decision. The BI tool is the delivery mechanism, not the point.
You don't always have the luxury of picking a fresh dataset. Sometimes you're working with what you have, or you've already started a project on a common dataset. Here's how to reframe existing work so it reads as real analysis rather than a tutorial.
Narrow the question. Instead of "analyze Superstore sales," pick one specific angle: "Which sub-categories had the highest return rate in the West region, and does order size predict returns?" That question is answerable from the Superstore data and reads as category management work.
Add a second data source. Joining the Superstore data to publicly available economic data by region, or to US population data by state, adds a layer of analysis that the tutorial never included. Suddenly the project isn't just exploring the dataset — it's answering something about the relationship between the dataset and the real world.
Extend the timeline or scope beyond the tutorial. If the tutorial covered 1 year of data, pull in 3 years and make the trend question central. If the tutorial filtered to one region, build a comparison across all regions with a specific hypothesis about why they differ.
Focus on a finding, not a display. Take the dashboard you already built and rewrite every chart title to state the finding, not the variable. "Sales by Region" becomes "Southeast Trails All Regions by 22% Despite Higher Traffic." Same chart, different intent, different impression.
The Analyst Hive program walks through the full project build sequence in Month 1 — which datasets to use, how to frame the question, and how to structure the output so it reads as real analytical work rather than coursework. The goal is 3 projects you can defend in any interview, not a checklist of completed exercises.
Even a strong project can underperform on a resume if it's framed badly. The most common mistake: listing the project as "Power BI Dashboard — Sales Analysis" with a link and nothing else.
That tells a reviewer the tool and the dataset type. It doesn't tell them what question you answered, what you found, or why they should click the link. They won't.
A stronger format treats each project like a result, not a deliverable:
Two to 3 lines per project is enough. The goal is a reviewer who reads the description and wants to see the work — not one who skips the link because the description didn't earn the click.
Where do I find datasets that aren't tutorial datasets?
Government open data portals (data.gov, your state or city's open data portal), sports reference sites (Baseball Reference, Basketball Reference, FBref), financial disclosures (SEC EDGAR, NCAA financials), healthcare data (CMS, CDC), and industry-specific sources (Zillow Research, BLS, Census Bureau) all have quality data that most candidates aren't using. Kaggle also has non-tutorial datasets if you avoid the competition datasets that appear in every tutorial. The key is picking a domain you can speak to, so the question you ask of the data feels natural rather than forced.
How long should a portfolio project take to build?
A SQL project — 2 to 4 queries with a README — should take 1 to 2 weeks of part-time work if you know the SQL concepts already. A BI dashboard project should take 2 to 3 weeks: 1 week on data prep and modeling, 1 to 2 weeks on the dashboard build and polish. If a project is taking significantly longer, the question is probably too broad or the dataset is too messy. Narrow the question and use cleaner source data.
Should my SQL project be multiple queries or one long query?
Multiple queries is usually cleaner and more readable. A SQL portfolio project works well as a set of 3 to 5 queries, each answering a sub-question that builds toward the main finding. Use CTEs to keep each query readable. Avoid nesting more than 1 level deep without a CTE to name the intermediate result. A reviewer who has to untangle 4 levels of nested subqueries to understand what you were doing forms a negative impression even if the output is correct.
Do I need to clean the data before building the project?
Yes, and documenting that you did is part of the portfolio signal. Analysts work with messy data. A candidate who can describe how they handled nulls, standardized inconsistent values, or restructured a flat file into a proper data model is demonstrating a real skill. Include a brief note in your README or dashboard methodology section about what the raw data looked like and what you did to prepare it. That context elevates the project.
Is a Python analysis project a good portfolio piece for analyst roles?
Yes, if Python is listed in the roles you're targeting. A pandas analysis in a well-documented Jupyter notebook — with a clear question, clean code, and a readable output — reads as real analytical work. Host it on GitHub. Include a README that explains the question and the findings. The same rules apply as for SQL and BI projects: question first, real dataset, specific finding. The tool is secondary.
What if I don't have time to build a new project from scratch?
Upgrade what you have. Pick your strongest existing project and apply the question-first reframe: rewrite chart titles to state findings, add a second data source to one of your SQL queries, add a filter or slicer to a static dashboard, rebuild the resume framing to lead with the question and the finding rather than the tool and the dataset. An upgraded existing project often reads significantly better than a rushed new one.
If you want a structured approach to building projects that read as real analytical work — with the question, the dataset, the model, and the resume framing all worked out — the Analyst Hive program covers exactly this in Month 1. Daily tasks, 3 projects built to a standard that gets interviews.