Selected work

Projects where data turned into decisions.

Analytics projects across sales, retail, social housing, financial fraud and pricing. Each one moves from a business question, through data modelling and analysis, to a dashboard or model a stakeholder could actually use.

Case study 01 · Sales performance analytics

Sales Performance Dashboard (Excel)

ExcelPivotTablesDashboard designData cleaning
Excel dashboard

Problem

A company's sales data sat in a flat, 12-column table — no clear view of which months, products, segments, factories, regions or suppliers were actually driving revenue, profit and cost.

Approach

  • Cleaned and structured the 12-column dataset ready for analysis.
  • Built PivotTables around the core business metrics — revenue, profit, unit economics and cost.
  • Designed a single-screen Excel dashboard with charts (and a separate table to drive the regional area chart, which PivotTables couldn't produce directly).

Result

A dashboard that turns a raw export into a clear read on where revenue and profit come from — and where the company is leaking margin.

32%Profit share — GreenTech (top)
$23.7KRevenue — South Australia (top)
>$60Unit cost — Factory C (highest)
12Data dimensions analysed

Insights & recommendations

Revenue trend & segments

Monthly revenue is trending down overall, with February the strongest month — a signal to act before it erodes further. Enterprise drives the most revenue and medium firms the least, so the growth lever is targeted marketing to medium and small businesses.

Product & unit economics

GreenTech leads on both revenue and profit (32%) while SolarMax trails (17%) — push GreenTech but keep the portfolio balanced. SolarMax and RenewTab have the highest profit per unit; lowering SolarMax's price could lift volume. EcoWidget has the weakest unit economics — improve its efficiency and marketing, or consider discontinuing it.

Factories & regions

Factory C is the costliest at >$60/unit vs Factory B at $50 — worth investigating C's day-to-day operations for discrepancies. South Australia leads regional revenue at $23,700 while ACT lags at $2,500, pointing to marketing upside in under-served territories.

Suppliers

Supplier X sold the most with a favourable unit cost; Supplier Y sold least but had the lowest production cost; Supplier Z was most expensive at >$1,200/unit on average. Recommendation: shift purchasing toward Supplier X and reduce reliance on Supplier Z.

Reflection

I built the PivotTables and visuals around profitability, revenue and cost so each one maps to a clear business metric. My biggest challenge was the revenue-by-area chart — it couldn't be produced from a PivotTable, so I created a separate table from the same dataset to feed it. The visual still isn't as clear as I'd like, but the project taught me a lot about structuring data for reporting and translating it into recommendations.

Case study 02 · Retail sales & customer analytics

O&F Retail Sales & Customer Analytics Storyboard

Office & Furniture retailerJul – Nov 2025

TableauExcelData storytelling
Tableau storyboard

Problem

An Office & Furniture (O&F) retailer had four years of orders (2019–2022) but no connected view of sales, profitability, customers and shipping — so it was hard to see where a thin margin was being eroded.

Approach

  • Cleaned four years of O&F sales in Excel, then built 3 Tableau dashboards and an executive storyboard: Sales Performance, Customer Segmentation, and Fulfilment & Shipping.
  • Analysed the sales & profit trend, category/sub-category performance and revenue by region, using Tableau calculated fields and dashboard actions.
  • Segmented customers into 3 types and 4 clusters to isolate the most profitable groups, and modelled a 5-day shipping SLA to quantify breach cost.

Result

  • Surfaced $3.59M in sales at a 6.97% average margin and $436K profit — with Oceania the top region ($1.1M+) and Southeast Asia loss-making.
  • Quantified $200K+ profit lost to a 5-day SLA breach (+$636.6K profitable vs −$200.5K loss-making) and flagged Brisbane (3.5-day vs 4-day average) for operational focus.
$3.59MTotal sales
6.97%Average margin
$436KTotal profit
3Tableau dashboards

Insights & recommendations

Insights

  • Profitability is real but uneven: $3.5M+ profit at a ~6.97% margin, yet 2021 profit stayed flat while sales rose — a hidden-cost signature. Technology was broadly profitable and furniture chairs sold well, but tables made heavy losses; Oceania led at $1.1M+ while South-East Asia sold well but lost money.
  • Customers differ sharply in value: ordinary consumers were the most profitable segment and home-office the least, and clustering isolated one clearly most-profitable group (Cluster 4) as a precise retention target.
  • Operations quietly erode profit: average shipping was ~4 days (inside the 5-day SLA), yet SLA breaches still cost $200k+, with order volume heavily concentrated in Brisbane.

Recommendations

  • Grow the Technology category and innovate within Furniture, while investigating South-East Asia's losses and auditing 2021's accounts for the hidden costs before they compound.
  • Focus retention on the profitable consumer segment and Cluster 4, and investigate why home-office customers underperform rather than spending equally across all segments.
  • Protect margin operationally by holding deliveries within the 5-day SLA (e.g. alternative routing) and diversifying beyond Brisbane, using same-day delivery to lift satisfaction in other cities.

Reflection

Anchoring three dashboards to three explicit business goals and tying them into one storyboard taught me to think in narrative rather than isolated charts — the analysis only became persuasive once the profitability, customer and operational views connected into a single argument. Digging past the headline profit to the flat-profit / rising-sales pattern in 2021 showed me the value of asking why of a number instead of just reporting it, and translating that into a decision a business could act on.

Case study 03 · Social & industry analytics

Co-operative Industry Opportunity Analysis

Mercury Co-operative · PACE ProjectFeb – Jun 2026

TableauSQLExcelResearch
Interactive BI

Problem

Mercury Co-operative's national Co-op Map was static — hard for decision-makers to explore where co-op density, housing need and opportunity actually concentrated.

Approach

  • Cleaned and modelled data on 120 co-operatives from Mercury, ABS Census & SEIFA 2021 and CPI (SQL prep, Excel cleaning) to compute co-op density per 100,000 people by state and industry.
  • Built composite housing-need (1–10) and mismatch-severity scores across 8 states and 5 vulnerable groups over a 70-year (1955–2024) timeline.
  • Extended the Co-op Map into an interactive Tableau environment (maps, filters and parameters) for 3 decision-maker personas, under data-quality and visual-ethics safeguards.

Result

  • Mapped 39 housing co-ops and the 6 co-ops within 15km of campus, giving each persona a self-serve path from industry gaps to housing need.
  • Flagged QLD, SA and TAS as priority markets based on density and severity scoring.
120Co-ops modelled
3Decision-maker personas
1955–202470-year timeline
QLD·SA·TASPriority markets flagged

Insights & recommendations

Insights

  • Co-op presence is highly concentrated: NSW dominates most sectors while several states show clear sector gaps — revealing the highest-priority regions for Mercury to recruit members.
  • A real employability signal: 6 co-ops within 15km of the Macquarie University campus, with industries mapping cleanly onto MQ graduate disciplines, making targeted placement partnerships feasible.
  • Housing supply is badly mismatched to need: QLD and VIC carry large stressed populations with minimal targeted co-op supply, and the timeline shows co-op formation collapsing after the 1996 funding cuts even as the CPI rent index climbed from ~62 to 125+.

Recommendations

  • Refresh the underlying ABS, SEIFA and labour-force datasets annually — the Tableau workbook updates automatically when source files are replaced.
  • Expand the co-operative dataset beyond NSW, where the sample is thin, to strengthen cross-state comparisons.
  • Use the state-mismatch and 40-year timeline visuals directly as advocacy evidence in policy submissions and funding applications.
  • Connect the dashboards to suburb-level rental data and ASIC / ABS APIs to drill below state level and automate updates.

Reflection

Designing for three distinct personas forced me to treat every chart as a decision aid rather than decoration — each visual had to answer a specific question a specific stakeholder would ask. The most valuable part was the ethics work: choosing a single-hue blue scale instead of red-green so no state reads as simply 'good' or 'bad', fixing the CPI axis so the rent line couldn't overstate a correlation, labelling absent sectors as 'opportunity gaps' rather than failures, and keeping vulnerable-group data at aggregate state level with street addresses removed. I learned that responsible visualisation is a design constraint from the start, not a disclaimer added at the end.

Case study 04 · Risk & anomaly analytics

Credit Card Fraud Detection Analysis

State of Delaware · public datasetFeb – Jun 2025

Power BIPythonPandasDAX
Reporting

Problem

The State of Delaware's public "Credit Card Spend by Merchant" dataset held 1.6M+ transactions where anomalous, unusually high spend and redundant records were buried in volume.

Approach

  • Ran EDA in Python (Pandas) to detect anomalies and redundant records, and to separate genuine fraud from seasonal spending.
  • Classified transactions and profiled spend by merchant, category, department and time.
  • Built an interactive Power BI report with DAX measures, slicers and data storytelling.

Result

  • Flagged ~7.36% of transactions as fraudulent or unusually high, with a June spike to ~9% (May–July) and holiday spending (Nov–Jan) separated from genuine fraud.
  • Named Grainger the top fraudulent merchant and surfaced the Department of Corrections and telecom firms as top-spend accounts across $680M+.
1.6M+Transactions analysed
7.36%Flagged as fraud/high-risk
~9%June fraud peak
$680M+Total spend explored

Insights & recommendations

Insights

  • About 7.36% of transactions were flagged as unusually high or potentially fraudulent, giving a clear anomaly baseline against the 92.64% classified as normal.
  • Fraud risk is seasonal and merchant-specific: March–April carried the highest volumes and June saw anomalies spike to ~9%, while a handful of merchants and telecom vendors accounted for a disproportionate share of high-value activity.
  • Spend concentration is itself a risk signal — the Department of Corrections was by far the largest spender and the biggest vendors were dominated by telecom providers, both warranting closer review.

Recommendations

  • Focus monitoring on the May–July window, and June in particular, where anomalous transactions climb most sharply.
  • Prioritise audits of the highest-spending department (Corrections) and the top telecom merchants, where concentration increases both exposure and fraud potential.
  • Treat holiday-period spikes (Nov–Jan) with context — higher spend there is expected and shouldn't be auto-flagged, reducing false positives.

Reflection

This project sharpened my judgement about the difference between an anomaly and a genuine problem: the same 'unusually high' flag meant fraud risk in June but simply reflected expected seasonal spending in December. Building slicers for season, month, day of week and merchant taught me to design a dashboard that lets an investigator interrogate the data rather than just read a headline number. Working with a 1.6-million-row public dataset also reinforced practical BI skills — aggregation, ranking merchants by volume and value, and presenting a defensible 7.36% anomaly rate an executive could act on.

Case study 05 · Predictive modelling

Used-Car Price Prediction

Kaggle CompetitionJul – Nov 2025

PythonPandasscikit-learnXGBoost
Machine learning

Problem

A Kaggle competition asked for accurate listed-price predictions across 8,000 training vehicles with mixed categorical and numeric features.

Approach

  • Built an end-to-end regression pipeline in Python (Pandas, NumPy, scikit-learn).
  • Engineered features via imputation, one-hot encoding, scaling and regex-based extraction.
  • Trained and hyperparameter-tuned Ridge, HistGradientBoosting and XGBoost with GridSearchCV cross-validation.

Result

  • Selected XGBoost as best at 11.3% CV MAPE — less than half the 25.4% linear baseline error.
  • Produced a repeatable pipeline that can be re-run on new listings without manual rework.
11.3%CV MAPE (XGBoost)
25.4%Linear baseline error
8,000Training vehicles
3Models compared

Insights & recommendations

Insights

  • Price is driven by performance and usage: horsepower (0.63), torque (0.62) and a power-to-weight proxy (0.60) were the features most correlated with price, ahead of mileage (0.50) and car age (0.47).
  • The target is heavily right-skewed — most cars cluster between $18k and $37k (median $25.6k) with a long tail to ~$315k — making MAPE a more honest error metric than absolute error.
  • Non-linear models captured this far better than a linear baseline: XGBoost 11.3% CV MAPE and HistGradientBoosting 11.5%, versus 25.4% for Ridge — less than half the baseline error.

Recommendations

  • Adopt XGBoost as the production model; it won on cross-validated MAPE and generalised best on the leaderboard submissions.
  • Invest further feature engineering in the performance and usage variables, since they carry most of the predictive signal.
  • Model log-price or segment the luxury long tail separately, so a few very high-value cars don't distort error on the mass market.

Reflection

This project taught me that disciplined preparation matters more than model choice — most of the effort went into imputation, encoding, scaling and regex extraction before any model was trained, and that groundwork let three very different algorithms run on the same clean feature set. Building the tuning loop with GridSearchCV and comparing models on a single, interpretable metric (MAPE) made the final decision defensible rather than arbitrary. Seeing XGBoost halve the linear baseline was a concrete lesson in when tree-based methods are worth the added complexity; next time I would address the skewed price distribution earlier and guard harder against overfitting on the luxury tail.

Want the detail behind a project?

Happy to walk through the data, the decisions and what I'd do differently next time.

Get in touch