Examples — overview
All example datasets and project templates that ship with DMAIC.io. Every entry opens directly in the matching module via deeplink.
89 examples across 54 modules
Define
5 Why: Delivery Time > 30 min
5-Why chain from symptom 'pizza arrives late' down to root cause 'order system without logistics function'.
DMAIC Project Plan Pizza Delivery
32 events across 7 weeks — from kickoff to project closure, with a gate review per phase. Dates are remapped to the current week (Monday anchor) on load so the plan always centers on today.
Pizza Delivery (SIPOC)
Suppliers, Inputs, Process, Outputs, Customers for the Bella Margherita delivery process — 8 process steps from order intake to handover.
Pizza Delivery Project Plan (todo list)
15 tasks for the pizza delivery project across the DMAIC cycle — from kickoff and SIPOC through Ishikawa and hypothesis test to the reaction plan and lessons learned. Status, owner and target dates spread realistically: three done, three in progress, eight open, one blocked (waiting on a sponsor decision). Items with the calendar flag enabled appear in the DMAIC calendar as soon as that module is added to the Define tile.
Pizza Delivery Time (Project Charter)
Teaching case 'Bella Margherita': delivery times too long, cold-pizza complaints rising. Pre-filled with problem statement, three measurable goals, and the project-team org chart.
RACI Pizza Delivery
RACI matrix for 8 activities × 6 roles — who is R(esponsible), A(ccountable), C(onsulted), I(nformed)?
Stakeholder Analysis Pizza Project
7 stakeholders from CEO to accounting with power/interest ratings and support attitude (supporter / neutral / resistor).
VoC → CTQ Pizza Delivery
3 Voice-of-Customer quotes (cold pizza, long delivery, address issues) translated into CTQs with measurable requirements.
Measure
Bimodal Process
Mixture of two normal distributions around 9.93 and 10.07 mm. Violates the normality assumption of Cpk.
Bolt Diameter (capable)
100 normally distributed measurements around μ ≈ 10.00 mm with σ ≈ 0.05 mm. Textbook capable-process example.
MSA Type 1: capable gage
25 repeated measurements of a 10.0 mm reference part (LSL 9.9 / USL 10.1, T = 0.2 mm). Very small spread, negligible bias → Cg ≈ 3.8 and Cgk ≈ 3.8 well above the 1.33 threshold. Textbook capable-gage case.
MSA Type 1: incapable gage
25 repeated measurements with large spread and noticeable bias. Cg ≈ 0.6 and Cgk ≈ 0.3 — gage clearly fails the capability threshold. Textbook case for a measurement system that must be rejected.
MSA Type 1: marginal gage
25 repeated measurements with moderate spread and slight positive bias. Cg ≈ 1.5 just above the 1.33 threshold, Cgk ≈ 1.0 just below — bias pulls capability down. Classic case for the role of Cgk alongside Cg.
MSA Type 2: acceptable Gage R&R
Crossed ANOVA study: 10 parts × 3 operators × 3 replicates (90 measurements). Small repeatability and reproducibility variation, dominant part variance → %GRR ≈ 6 % of study variation. Textbook acceptable measurement system.
MSA Type 2: problematic Gage R&R
Crossed ANOVA study with large repeatability and reproducibility variation versus small part variance. 10 parts × 3 operators × 3 replicates → %GRR in the 40–50 % range. Textbook case where the measurement system masks part-to-part variation.
Process Map Pizza Delivery
8 process steps with value-type classification (VA / BNVA) and input types (parameter x / noise n). Order → kitchen → bake → pack → deliver → handover.
Shifted Process
Mean off-center (μ ≈ 10.08 mm) → low Cpk despite small dispersion.
Analyze
Bolt diameter with outliers (Grubbs)
30 bolt diameters from N(10.000, 0.050) with two seeded outliers at 9.745 and 10.260 (≈ ±5σ). Grubbs (two-sided), generalized ESD and Tukey IQR should flag both.
C&E Matrix Pizza Delivery
Inputs (routing, oven, traffic, …) versus outputs (delivery time, temperature, complaint rate) — weighted 0–9 ratings of the dependencies.
Compare means of two machines (2-sample)
Compare cycle time of two machines: difference Δ = 0.5 min at σ ≈ 0.8 min. Two-sided t-test, α = 0.05, power = 80 %. Cohen's d ≈ 0.63 — returns the n per group.
Compare variances of two suppliers (2-sample)
Compare scatter of two suppliers: suspected doubled variance (σ²₁/σ²₂ = 2.0). Two-sided F-test, α = 0.05, power = 80 %. Returns the sample size per group.
Complaints by branch (5-sample ANOVA, unbalanced)
Complaints per day across five branches A–E, sample sizes 8/10/12/15/9. Branch D has a clearly elevated mean (6.2 vs. ~5.0) — post-hoc comparisons separate D from the rest.
Delivery time by driver (k-sample ANOVA)
Delivery times (min) for three drivers A/B/C, 15 each. True means 22/25/28 min at σ ≈ 3 — Cohen's f ≈ 0.82, strong ANOVA signal.
Detect variance reduction (1-sample)
After a process improvement, scatter should have dropped from σ₀ = 1.5 to σ₁ = 1.0. One-sided chi-square test, α = 0.05, power = 80 %. Returns the n needed to confirm the improvement statistically.
Fit exponential waiting times
80 waiting times from Exponential(μ=15 min) — memoryless distribution. Exponential and gamma rank highest.
Fit lognormal cycle times
80 cycle times from Lognormal(μ_log=ln 5, σ_log=0.6) — strongly right-skewed. Distribution-fit should rank lognormal first.
Fit normal distribution
80 values from N(μ=10, σ=2) — clean textbook normal distribution. Shapiro-Wilk / Anderson-Darling should not reject H₀.
Fit Weibull lifetimes
80 component lifetimes from Weibull(k=2.5, λ=100 h). Classic wear-out case (β > 1); Weibull should rank as best fit.
FMEA Pizza Delivery
4 risks along the delivery process (address entry, oven overload, thermal packaging, traffic) with S/O/D ratings and 2 dated actions each (one already completed, one planned) — burndown shows plan and actual lines.
Internal sales order processing
8 steps across 5 roles (customer, sales, dispatch, warehouse, accounting). Processing ≈ 53 min, waiting time dominates with ca. 2 d 13 h 30 min — PCE below 3 %. Showcases standard issues: inbox idle time, handover waiting, customer-side delays.
Ishikawa: Delivery Time Too Long
Full 6-M example: 6 hypotheses across Man, Method, Machine, Material, Environment, Measurement — rated by three experts — plus 7 supporting facts and 6 experiments (with status, schedule, cost, linked to hypotheses).
Mean shift from target (1-sample)
Check whether the process mean drifts by Δ = 0.3 from target (σ ≈ 0.5). Two-sided t-test, α = 0.05, power = 80 %. Cohen's d = 0.6 (medium effect).
Pizza Delivery Time (correlation matrix)
Same 30 pizza deliveries as the regression dataset — here as a correlation matrix over distance, traffic level and delivery time. Highlights the strong pairwise Pearson correlations.
Project selection criteria
Pre-filled criteria for project selection in the pizza-delivery process: delivery time, pizza quality, price, employee satisfaction, complaint rate. Pairwise comparisons are run inside the module.
Vacation request (classic approval flow)
9 steps across 4 roles, paper-based workflow from application to system booking. Processing ≈ 22 min, Lead Time > 4 workdays — PCE < 1 %. Classic office lever: one huge idle block in the manager's inbox.
Improve
Batch Complaints (NegBin GLM)
30 batches with batch size and supplier as predictors, complaints as response. Overdispersed count data — negative binomial regression additionally estimates the dispersion parameter θ.
D-optimal augment: on existing data set
Three continuous factors, quadratic model, 12 runs total — six already measured. The example creates a data sheet with the existing measurements and augments it with six additional D-optimal runs.
D-optimal mixed: 3 cont. + 1 categorical
Three continuous factors plus one three-level categorical factor (material). Mixed-level optimal designs are a textbook D-optimal use case — classic schemes cannot handle this.
D-optimal RSM: 3 factors, quadratic
Response-surface model with three continuous factors and a full quadratic polynomial in 12 runs. Classic D-optimal use case.
D-optimal screening: 5 factors, linear + 2FI
Five continuous factors with main effects and two-factor interactions in 18 runs. D-optimal alternative to Plackett-Burman when 2FI must stay in the model.
Defects per Shift (Poisson GLM)
40 shifts with feed rate (mm/s) and tool wear (h) as predictors, defect counts as response. Poisson model with log(λ) = -1 + 0.02·feed + 0.03·wear.
Fuel Consumption with Outliers
30 vehicle measurements speed → fuel consumption. Two deliberate outliers (rows 6 and 23) drop R² compared to the clean linear fit. Textbook case for residual diagnostics and outlier detection.
Motor Trials (multicollinearity)
30 motor trials with rpm and torque as predictors, power as response. Torque ≈ 0.04·rpm — predictors are highly collinear, VIF spikes. Textbook multicollinearity case.
Pizza Delivery Time (multiple linear regression)
30 pizza deliveries with distance (km), traffic level (1–5) and measured delivery time (min). Linear model Time ≈ 5 + 2.5·distance + 1.2·traffic — textbook dataset for multiple linear regression with high R².
Solder Inspection (logistic GLM)
50 PCBs with solder temperature (220–255 °C) and pressure (2.0–3.5 bar) as predictors, defect (0/1) as response. Textbook case for binary logistic regression without a trials column.
Surface Defects (Poisson GLM)
30 production lines with speed (parts/min) and line age (years) as predictors, surface defect counts as response. Poisson model for predicting the expected defect rate.
TRIZ 9 Windows Pizza Delivery
3×3 matrix Sub-/System/Super-system × Past/Present/Future — applied to the pizza delivery process.
TRIZ Contradiction: Speed vs. Productivity
Classic TRIZ contradiction from the pizza case: deliver faster without sacrificing productivity. Improving = 9 (Speed), worsening = 39 (Productivity).
TRIZ Evolution Trends: Passenger-Car Headlight
Eight classical evolution trajectories applied to a passenger-car headlight — from twin-headlight to adaptive LED matrix and beyond.
TRIZ IFR: Belt Conveyor
Three IFR levels for a jamming belt-conveyor transfer chute, with an obstacle list as the task backlog.
TRIZ Physical Contradiction: Aircraft Wing
Classic textbook case: a wing must be both long (lift) and short (drag). Resolved via separation in time — swing wing.
TRIZ Resources Checklist: Belt Conveyor
Resources inventory across 6 categories × 3 system levels for the belt-conveyor case — companion to the IFR example.
TRIZ Substance-Field: Drill / Concrete
Su-Field model for an insufficient mechanical action (drill on concrete). Diagnosis: class 2 — pulsed mode (hammer drilling) selected.
Yield Trials (logistic GLM)
Eight temperature levels with 50 trials each. Binomial response (successes / trials) follows a logistic curve with inflection around 160 °C. Textbook case for binomial GLM with a trials column.
Control
Box-Cox I-MR chart: lognormal lifetimes
50 lognormally distributed lifetimes. A plain I-MR chart raises false alarms because of the right skew. After Box-Cox transformation (λ ≈ 0) the residuals become approximately normal.
c-Chart: defects per unit
30 inspection units of constant size. Mean defects ≈ 3, units 22–23 spike to ~8 — c-chart triggers.
EWMA chart: small drift
60 observations: mean stays at μ=10 for the first 30 points, then shifts to μ=10.1 (≈ 1σ). EWMA with λ=0.2 detects the drift; a Shewhart I-chart misses it.
Hotelling T²: bivariate correlation
40 bivariate observations with ρ ≈ 0.85. Observations 30 and 31 are multivariate outliers — plausible on each margin but joint deviates from the correlation ellipse.
I-MR Chart with Drift
50 individual measurements with a linear upward drift (+0.6 across the series). Nelson rule 3 (six points in a row trending up) should fire multiple times.
I-MR Chart with Shift
50 individual measurements with a step shift from μ = 10.00 to μ = 10.18 at observation 30. The Nelson rules (especially rule 1 and rule 5) should flag the shift.
Lessons Learned Pizza Pilot
4 lessons from the pilot: pre-route planning (success), double-wall thermal boxes (success), Monday standup (improvement), order intake bottleneck (problem).
p-Chart: variable sample sizes
30 inspection lots with variable sample size (80–120) and varying defect rate. Lots 18–20 at ~10 % — the p-chart should flag them as out-of-control.
Shaft Diameter (X̄-R, in control)
25 subgroups of 5 measurements each from a stable process (μ = 10.00 mm, σ = 0.05 mm). X̄-R chart shows no Nelson-rule violations.
Short-Run Z-MR chart
Three short setups with different means (10, 25, 7.5), 8 measurements each. After nominal centering they form one comparable process (σ ≈ 0.05).
t-Chart: days between rare events
25 intervals from Exponential(μ=35 days). At observations 16–18 the rate doubles (μ=12) — t-chart shows a visible shift.
Data & Tools
Box-Cox transformation: lognormal cycle times
80 strongly right-skewed lognormal cycle times. Box-Cox (optimal λ ≈ 0) makes the values approximately normal.
Complaint Cost Q2 (Pareto)
50 Q2 complaints with reason and individual cost. The Pareto analysis sums cost per reason and sorts descending — the 80 % line shows that 'dimensional deviation' alone accounts for most of the cost, even though it is not the most frequent reason. Textbook example: frequency ≠ importance.
Complaint reasons Q1 (donut)
248 customer complaints from Q1, split across six reasons. Donut variant (innerRadius = 0.5) with the total shown as the center label — useful when the total itself carries a key message.
Cycle Time Shift × Line (heatmap)
45 cycle-time measurements (in seconds) across three shifts (early/late/night) and three production lines (L1/L2/L3) — 5 readings per combination. The heatmap reveals two effects at once: a line effect (L3 ≈ 30 s slower than L1) and a shift effect (night ≈ 10 s slower than early). The (night, L3) cell at 82 s is the clear worst case — a perfect starting point for the Analyze phase.
Defects per Shift (bar chart)
60 defect events across three shifts (Early/Late/Night) and four defect types (scratches, dimensional, surface, other). The stacked bar chart shows at a glance that the early shift is dominated by dimensional defects while the night shift produces mostly surface defects — a classic anomaly for the Measure phase.
Delivery time by driver (boxplot)
Three boxes showing delivery-time distributions for drivers A/B/C. Same data as the ANOVA example — the boxplot visualizes the ANOVA effect.
Delivery time by driver (individual value plot)
Three point clouds of delivery time for drivers A/B/C with connected means — same data as the ANOVA and boxplot examples.
Fuel Consumption with Outliers (scatter plot)
Same 30 vehicle measurements as the regression dataset — scatter plot of fuel consumption over speed with stats panel. Two deliberate outliers (rows 6 and 23) are directly visible in the plot.
Inspection Log (all column types)
15-row inspection log that exercises all seven worksheet column types: text (inspection ID, shift), date, time, numeric (diameter in mm), currency (unit cost €), percent (scrap rate) and binary (OK 0/1). Includes a few null cells on purpose to show how missing values are handled.
Machine Fleet (bubble chart)
10 machines with cycle time, scrap rate and daily output. The bubble chart prioritises Six Sigma improvement projects: large bubbles in the upper right (high scrap at high output) are the most expensive problem cases.
Motor Trials (multicollinearity)
Same 30 motor trials as the regression dataset — two scatter plots of power over rpm and power over torque, with stats panel. Visualises the strong predictor correlation before fitting a regression.
Orders with Formula Columns
10 order rows with quantity, unit price (€), calculated total (=qty·price), discount (%) and final price (=total·(1-discount)). Demonstrates cell-reference formulas and currency/percent columns.
Pizza Delivery Time (scatter plot)
Same 30 pizza deliveries as the regression dataset — scatter plot of delivery time over distance with stats panel. Useful for visual exploration before fitting a regression.
Production Line Defects (bubble chart)
10 production lines with defect rate, repair cost per defect, and monthly defect count. The bubble chart shows the classic frequency ↔ unit-cost trade-off; bubble size reveals which lines carry the largest absolute loss potential — an ideal entry point for DMAIC project selection.
Production Q1–Q3 (multi-sheet demo)
Three sheets (Q1/Q2/Q3) with production data of 10 batches each. Demonstrates the multi-sheet feature of the worksheet — switch via tabs at the bottom.
Project Portfolio (bubble chart)
12 Six Sigma project candidates with effort (person-days), annual benefit (k€) and risk score, grouped by area (Production vs. Quality). The bubble chart prioritises project selection while showing how the two areas position themselves in the effort/benefit/risk landscape: small bubbles in the upper left are quick wins; large bubbles in the lower right should be deferred.
Q-Q plot: comparing two shifts
Two shifts of 40 measurements each — shift A ≈ N(μ=9.5, σ=1.5), shift B ≈ N(μ=10.8, σ=2.2). The probability plot shows two lines with different location and slope.
Q-Q plot: measurement resolution problem
80 bolt diameters measured with a gauge that is too coarse (resolution 0.1 mm versus process σ ≈ 0.15 mm). The probability plot shows clearly visible vertical stripes — a classic indicator of insufficient measurement resolution.
Q-Q plot: Weibull lifetimes
80 Weibull(k=2.5) lifetimes on the normal probability plot — visible curvature in the tails.
Run-Chart with shift
I-MR series: observations 1–29 stable at μ=10, from obs. 30 shift to μ=10.18. Western Electric / Nelson rules flag the shift.
Scrap by defect type
395 scrap parts from a production line, split across six mutually exclusive defect types. Textbook pie-chart use case: share of a few categories of a whole. Full-circle variant (innerRadius = 0).
Supplier Approval per Plant (mosaic)
60 approval decisions across three plants and three status levels (OK/conditional/rejected). The mosaic plot reveals two effects at once: the volume differences (column widths — plant A has twice as many decisions as plant B, three times more than plant C) and the quality differences (segment heights — plant C has clearly more rejections per plant).
Supplier Scorecard (bubble chart)
12 suppliers with price index, on-time delivery rate and annual order volume. The bubble chart visualises the classic price ↔ reliability trade-off; bubble size reveals which suppliers carry the largest strategic weight.
Yield Optimum (response surface)
Quadratic model for yield = f(temperature, pressure) with optimum at 180 °C and 3 bar. Shows a classic contour map with a pronounced maximum.