Ian Klosowicz

Retail and e-commerce is one of the most common industries for entry-level analyst roles, and one of the worst industries to build a generic portfolio for. The Superstore dataset has appeared in so many tutorials that hiring managers at retail companies recognize it instantly and mentally file your project under "coursework." The fix isn't to avoid retail — it's to use real retail data with a real analytical question.
This post covers 8 specific project ideas using publicly available retail and e-commerce data, what each one demonstrates technically, and why each one reads as work a retail analytics team would actually do rather than a tutorial exercise.
Retail analytics covers a wide range of work depending on whether you're at a brick-and-mortar retailer, a pure-play e-commerce company, or a marketplace. The analytical problems that appear most consistently:
A portfolio project that maps to any of these makes it immediately clear that you've thought about what retail analysts actually do, not just what data looks like in a tutorial. That framing changes how the project reads before the hiring manager has looked at a single chart.
The Superstore dataset isn't the only option. Most of the data sources below are either free or freely accessible and produce better portfolio projects than any tutorial dataset.
US Census Bureau Retail Trade Survey publishes monthly retail sales by category (electronics, clothing, food, sporting goods, etc.) at the national level going back decades. Access at census.gov/retail. Good for macro trend analysis and category comparison projects.
UC Irvine Machine Learning Repository has several real e-commerce transaction datasets that aren't tutorial staples. The Online Retail II dataset covers 2 years of transactions from a UK-based online retailer with customer IDs, product descriptions, quantities, prices, and dates — a natural fit for customer analytics and cohort work.
Bureau of Labor Statistics Consumer Price Index publishes retail price indices by category over time. Combined with sales data, it enables real price elasticity analysis. Access at bls.gov/cpi.
Kaggle non-tutorial retail datasets include transaction data from real companies published for research purposes. The Brazilian E-Commerce Public Dataset by Olist (100,000+ orders with product, seller, customer, and review data across multiple tables) is structured as a proper relational dataset and is not a common tutorial choice. Search Kaggle for "olist" or "brazilian ecommerce."
Walmart and Target open datasets surface occasionally through research partnerships and data science competitions. Check Kaggle and the UC Irvine repository for recent additions.
Google Trends provides search interest data by term and geography that can be used as a proxy for consumer demand and combined with sales data for demand forecasting projects.
USDA Economic Research Service publishes food retail data including grocery store counts, food access metrics, and food expenditure data by geography. Useful for food retail and grocery analytics projects.
1. Customer cohort retention analysis
Question: Of customers who made their first purchase in a given month, what percentage return for a second purchase within 30, 60, and 90 days, and how does that retention rate vary by product category or acquisition channel?
Data: UC Irvine Online Retail II dataset or the Olist Brazilian E-Commerce dataset. Both have customer IDs, purchase dates, and product information needed for cohort analysis.
What it demonstrates: cohort construction using window functions (first purchase date per customer), date arithmetic, retention rate calculation, a dashboard with cohort heatmap or retention curve by acquisition month. This is one of the most common analyses in e-commerce and one of the most technically impressive things an entry-level candidate can show.
Why it reads as real: every e-commerce company tracks cohort retention. It's the primary lens for understanding whether the business is growing sustainably or churning through one-time buyers. A candidate who can build this without being asked has demonstrated awareness of the metric that matters most.
2. Product category performance and margin analysis
Question: Which product categories drive the most revenue versus the most margin, and which categories are growing fastest relative to their current size?
Data: Census Bureau Monthly Retail Trade Survey combined with BLS Consumer Price Index data to build real-dollar (inflation-adjusted) category comparisons over time.
What it demonstrates: multi-source join on category and date, inflation adjustment calculation, growth rate analysis using window functions, a dashboard with category comparison matrix (revenue vs. margin quadrant or growth vs. size bubble chart). Inflation adjustment is a real skill that most tutorial projects skip entirely.
Why it reads as real: merchandising teams at every retailer run category performance reviews. The specific question — revenue vs. margin vs. growth — is exactly the framework a category manager uses to prioritize assortment decisions.
3. RFM customer segmentation
Question: Based on recency, frequency, and monetary value of purchases, how do customers segment, and what does each segment's behavior look like over the next 6 months?
Data: Olist dataset or Online Retail II. Both have enough transaction history to build meaningful RFM segments.
What it demonstrates: RFM score calculation using window functions and CASE WHEN, customer segmentation logic in SQL, segment-level behavioral analysis, a dashboard showing segment size, average order value, and purchase frequency by segment. RFM is a standard retail analytics framework that signals real domain knowledge.
Why it reads as real: CRM and loyalty teams at retailers use RFM segmentation to target promotions, personalize communications, and identify at-risk customers before they churn. Knowing the framework and being able to build it is a meaningful differentiator.
4. Promotional discount effectiveness
Question: Do products sold at a discount generate enough incremental volume to offset the margin impact, and does the answer vary by product category or discount depth?
Data: Olist dataset includes price and freight data per order item. Online Retail II has unit prices that vary across transactions, enabling discount vs. full-price comparison.
What it demonstrates: revenue vs. margin trade-off analysis, discount depth segmentation (0-10%, 10-20%, 20%+), volume comparison at different price points, SQL aggregations by discount bucket, a dashboard with promotional lift visualization. This is pricing analytics work that most entry-level candidates have never touched.
Why it reads as real: markdown and promotional effectiveness is a constant concern in retail. Buyers and pricing teams run exactly this analysis to determine whether a promotion paid off or just gave margin away.
5. Return rate analysis by product and channel
Question: Which product categories and fulfillment channels have the highest return rates, and is there a relationship between return rate and customer retention?
Data: Olist dataset includes order status (delivered, canceled, returned) and product category. The Online Retail II dataset has negative quantities that represent returns.
What it demonstrates: handling negative transactions and returns in SQL, return rate calculation by category and channel, correlation analysis between return rate and repeat purchase behavior, a dashboard with return rate trend over time and by segment. Handling returns correctly in SQL (negating them from revenue, not just filtering them out) signals real analytical maturity.
Why it reads as real: returns are a major cost center in e-commerce. Analytics teams actively analyze return patterns to identify high-return SKUs, size and fit issues, fulfillment problems, and fraud. This is a live operational concern at every online retailer.
6. Retail sales trend and seasonality analysis
Question: How do retail sales by category respond to seasonal patterns, and which categories show the strongest year-over-year growth after accounting for seasonality?
Data: Census Bureau Monthly Retail Trade data going back 10+ years across 12+ retail categories. Clean, structured, directly downloadable.
What it demonstrates: time series decomposition (identifying trend vs. seasonal component), year-over-year growth calculations, category comparison over multiple economic cycles, a line chart dashboard with category filter and period selector. Working with 10+ years of monthly data is a meaningful scale signal.
Why it reads as real: retail planning teams use seasonality analysis to set inventory levels, staffing plans, and promotional calendars. Any retailer with a planning function runs this kind of analysis annually.
7. E-commerce seller performance analysis
Question: On a marketplace platform, which seller characteristics (category, location, review score, shipping speed) are most strongly associated with high sales volume, and do the same factors predict customer satisfaction?
Data: Olist dataset is a marketplace with seller, product, order, review, and customer tables — 5 related tables with proper foreign keys. This is one of the cleanest multi-table relational datasets available for portfolio work.
What it demonstrates: 4 to 5 table SQL join across a proper star schema, seller-level aggregations, correlation between operational metrics (shipping speed) and outcome metrics (review score, repeat purchases), a dashboard with seller performance scorecard. The Olist data model is a genuine relational schema that demonstrates real data modeling skills.
Why it reads as real: marketplace analytics teams at platforms like Amazon, Etsy, or Walmart Marketplace run exactly this analysis to identify seller quality issues, inform seller support priorities, and improve buyer experience.
8. Geographic retail opportunity analysis
Question: Based on population density, income levels, and existing retail presence, which metropolitan areas are underserved by a specific retail category, and how does that gap correlate with consumer spending data?
Data: Census Bureau population and income data by metro area, Bureau of Labor Statistics Consumer Expenditure Survey for spending by category, USDA food access data or similar geographic retail density data.
What it demonstrates: multi-source geographic join (metro area as the key), gap analysis between spending potential and existing retail supply, ranking metros by opportunity score, a map-based or ranked dashboard with drill-down by category. This is a site selection and market expansion analysis that retailers run before opening new locations.
Why it reads as real: real estate and strategy teams at retailers use exactly this analysis to evaluate market entry decisions. For a candidate without retail experience, building this project demonstrates understanding of how retail growth decisions are made.
If you want a structured approach to picking the right project from this list, scoping the question tightly, and building it to a standard that gets interviews, the Analyst Hive program covers the full build sequence in Month 1.
The analytical vocabulary of retail is one of the most accessible of any industry — most people have been customers their whole lives and understand the business problems intuitively. Use that.
When you walk through a retail portfolio project, connect the finding to a business decision a real team would make. "Customers who receive their order within 2 days have a 40% higher 90-day retention rate" leads naturally to "which suggests the logistics investment in faster fulfillment has a measurable return on the customer side, not just the operational side." That connection from data to decision is what separates a strong walkthrough from a description of what the dashboard shows.
If you have retail or e-commerce work experience — even in a non-analytical role — connect the project to problems you observed. A candidate who worked in retail operations and built a return rate analysis because they saw how returns were handled on the floor is a genuinely different signal than a candidate who picked the dataset because it was available. Use that context.
Can I use the Superstore dataset if I do something original with it?
You can, but the dataset recognition still happens and some of the benefit-of-the-doubt you'd get with an unfamiliar dataset disappears. If you've already built something on Superstore and it's strong, reframe the question as specifically as possible and make sure the resume description leads with the finding, not the dataset name. For a new project, one of the datasets above is a better starting point.
Is the Olist dataset good for a SQL portfolio?
It's one of the best publicly available datasets for a SQL portfolio project specifically. It has 5 related tables with proper foreign keys, real transaction data across 100,000+ orders, and enough dimensions (product category, seller location, customer location, review score, shipping time) to write 3 to 5 interesting queries that each answer a different sub-question. The data is in Portuguese, but the column names are descriptive enough that it's easy to work with.
What's the difference between a retail analyst role and an e-commerce analyst role?
Retail analyst roles at brick-and-mortar companies tend to focus more on store-level performance, inventory, merchandising, and in-store operations. E-commerce analyst roles focus more on digital funnel metrics (conversion, cart abandonment, traffic attribution), customer behavior online, and fulfillment. Marketplace analyst roles at companies like Amazon or Etsy add seller-side analytics. The projects above cover all 3 contexts — pick the ones that map to your target role type.
How do I handle the Olist data being in Portuguese?
The product category names are in Portuguese, but the dataset includes an English translation table that joins on category name. Include that join in your SQL and use the English names in your dashboard. Document in your README that you used the translation table — it shows you read the data dictionary, which is a real analyst skill.
Which project is most likely to impress at an interview for a retail analytics role?
The cohort retention analysis (#1) and the RFM segmentation (#3) are the 2 that most consistently signal genuine retail analytics knowledge, because they use frameworks that retail teams actively use and that most entry-level candidates don't know exist. If you can explain what a cohort retention curve shows and why it matters for a subscription or repeat-purchase business, you've demonstrated more domain awareness than most candidates in the pool.
Do I need to build all 8 of these?
No. 1 to 2 strong projects in the retail space are enough. Pick the question you find most interesting and that best matches your target role, build it well, and be able to walk through it in depth. A single cohort retention analysis you can defend confidently beats 4 surface-level projects in any interview.
If you want a structured path through picking the right project, building it to a standard that gets read, and framing it on a resume that earns clicks, the Analyst Hive program covers all of it in Month 1. Daily tasks, built around getting hired.