Everything the lecture
never actually said
Rebuilt from the course pack: the syllabus, the one PPT that exists, every line of lecture code, and the real datasets. Every number on this page is computed from Mahakoshal Textiles' actual demand file, not invented for teaching.
Foundations
Operations analytics is the layer between raw operational data and a decision somebody has to make anyway: how much to produce, how much to stock, when to reorder, where to send it. You already do this every time you size a position with incomplete information about tomorrow. Same discomfort, physical goods instead of a price.
Strategic or tactical
The split decides how much analysis a decision deserves. Strategic means slow, expensive to reverse, worth weeks of modelling. Tactical means daily, cheap to adjust, needs something you can recompute every morning. Sort these six, taken from the course slides.
The rule everything else depends on
Before forecasting you split data into a training period and a test period. In ordinary machine learning you shuffle first. In time series that is a serious error, and you already know why from backtesting: shuffled training data can contain days that fall after your test days, so the model learns from a future it would never have had. It scores brilliantly, then collapses in production.
Drag the toggle and watch where the training points land.
Mahakoshal's series runs 417 days. Should the test set be a random 30 days scattered across it, or the most recent 30? And sort these: IIM Vizag opening a Hyderabad campus, versus the mess deciding tomorrow's rice quantity.
The most recent 30. A scattered sample would put training days after test days, which is leakage, the same sin as lookahead bias in a backtest. Campus is strategic, slow and expensive to reverse. Rice is tactical, decided daily and adjusted freely.
Naive and moving average
The naive forecast says tomorrow equals the last value actually observed. No maths, no learning. A moving average instead takes the mean of the last k days, which is the identical formula to the MA20 or MA50 line sitting on every chart you have ever opened on Kite. Kilograms of yarn instead of rupees per share.
Move the window. Watch the forecast line stiffen, and watch the four error numbers move with it. Every value recomputes live against Mahakoshal's real test month.
What those four words actually mean
Four full names, since none of them should be a mystery: Mean Absolute Error, Root Mean Squared Error, Mean Absolute Percentage Error, and Bias. All four start from the same first step: for a given day, error just means predicted minus actual. If you predicted 45 and 50 actually happened, your error is 45 minus 50, which is minus 5. Negative means you guessed too low, positive means you guessed too high. Nothing more than subtraction so far.
Let's build all four by hand, slowly, using one small example: a canteen sells a steady 50 cups of chai every day, and across five days you predicted 45, 55, 60, 50, 40.
| Day | Actual | Predicted | Error, predicted minus actual |
|---|---|---|---|
| 1 | 50 | 45 | -5 |
| 2 | 50 | 55 | +5 |
| 3 | 50 | 60 | +10 |
| 4 | 50 | 50 | 0 |
| 5 | 50 | 40 | -10 |
Mean Absolute Error, MAE. "Absolute" means: stop caring whether the error was positive or negative, just look at its size. Minus 5 and plus 5 both become 5.
Step 1, drop every sign: 5, 5, 10, 0, 10.
Step 2, add them up: 5 + 5 + 10 + 0 + 10 = 30.
Step 3, divide by how many days you added, 5: 30 ÷ 5 = 6.
MAE = 6. In plain words: on a typical day here, the prediction was off by 6 cups, could have been over or
under, this number alone does not say which.
Root Mean Squared Error, RMSE. Same goal as MAE, get rid of the sign and find a typical error size, but it gets rid of the sign by squaring instead of dropping it, and squaring has a side effect worth knowing about: a miss twice as large gets counted four times as harshly, not twice.
Step 1, square every error, a negative times a negative is positive so this also removes the sign:
25, 25, 100, 0, 100.
Step 2, average those squared values: (25+25+100+0+100) ÷ 5 = 250 ÷ 5 = 50.
Step 3, take the square root, this undoes the squaring and brings you back to plain cups instead of
cups-squared: √50 ≈ 7.1.
RMSE = 7.1, a little higher than MAE's 6. That gap is not noise, it is information: days 3 and 5 had the
biggest misses, size 10, and squaring punished those two harder than it punished the smaller ones. A bigger
gap between RMSE and MAE always means a few rough days are hiding inside an otherwise ordinary-looking
average.
Mean Absolute Percentage Error, MAPE. Same idea as MAE, average size of the miss, just measured as a percentage of that day's actual value instead of raw cups.
Step 1, for each day divide the absolute error by that day's actual, then multiply by 100 to get a
percent: day 1 is 5 ÷ 50 = 10%, day 2 is 5 ÷ 50 = 10%, day 3 is 10 ÷ 50 = 20%, day 4 is
0%, day 5 is 10 ÷ 50 = 20%.
Step 2, average those percentages: (10+10+20+0+20) ÷ 5 = 12%.
MAPE = 12%. Why bother, when MAE already told you 6 cups? Because "6 cups off" means something completely
different for a stall selling 50 cups a day than for a factory selling 50,000. MAPE lets you compare
forecast quality across completely different scales. Its one real weakness: it falls apart if actual demand
can be zero on some day, since you would be dividing by zero.
Bias. The one metric here that refuses to throw away the sign.
Step 1, keep every error exactly as it was, signs and all: -5, +5, +10, 0, -10.
Step 2, average them: (-5 + 5 + 10 + 0 - 10) ÷ 5 = 0 ÷ 5 = 0.
Bias = 0. Notice this is a genuinely different number from MAE's 6, not just a smaller version of it. MAE
said the typical miss was 6 cups. Bias says that over the five days, the over-guesses and under-guesses
happened to cancel out completely, so there is no lasting lean either way. That is the entire reason bias
gets its own column: MAE can never tell you which direction you are consistently wrong, because turning -5
and +5 both into 5 is exactly what erases the direction. Bias is the only one of the four that keeps it.
Now the instrument below has these exact five days already loaded. Change any prediction and watch all four numbers, and every step of the arithmetic above, recompute live.
| Day | Actual | You predicted | Error | Absolute | Squared | Percent |
|---|---|---|---|---|---|---|
| Sum | ||||||
Press the two buttons above and compare. Both sets can carry an identical MAE while their bias differs completely, because turning minus five and plus five both into five is exactly what erases direction. That is the whole reason bias exists as a separate number: it is the only one that tells you which way you are wrong, and therefore which failure is building up.
The scoreboard on real data
Computed from the actual file, last 30 days held out, rolling one step ahead forecasts.
| Method | MAE kg | RMSE kg | MAPE % | Bias kg |
|---|---|---|---|---|
| Naive | 29.7 | 36.0 | 1.23 | -0.2 |
| 3 day moving average | 22.5 | 27.2 | 0.94 | 0.2 |
| 7 day moving average | 20.1 | 26.0 | 0.84 | 2.1 |
Averaging a week instead of copying yesterday cuts average error by about a third. That works here because this stretch of demand is mostly noise around a stable level, nothing trending, nothing seasonal. Hold on to that caveat. The moment a series has real trend or seasonality, more smoothing stops being free, which is exactly where module 3 begins.
Worth noticing: the case narrative describes demand spiking past 2600 kg and dropping under 2300. The attached dataset never does either, its true range is 2322.9 to 2471.7 kg. Your professor's own code flags this and tells you to trust the data over the story.
Three most recent days were 2400, 2410, 2390 kg in that order. Give the naive forecast and the 3 day moving average. Then: a forecaster's errors are minus 5, minus 10, minus 5, 0, minus 10. What is their MAE and their bias, and which failure is building at the warehouse?
Naive is 2390, just copy the last value. The moving average is 2400 + 2410 + 2390 = 7200, divided by 3, so 2400. Not 2410, that is only the middle number, an average always sums then divides by the count.
The forecaster: absolute errors are 5, 10, 5, 0, 10, summing to 30, divided by 5 gives MAE of 6. Bias keeps the signs, so minus 30 divided by 5 gives minus 6. Every miss leans the same way, chronic under forecasting, so weavers run short. Nothing piles up in the warehouse, the shortage is on the floor.
Exponential smoothing
Naive and the moving average both have an odd habit: every day inside the window counts equally, and the moment a day falls outside the window it counts for nothing at all. Exponential smoothing, SES for short, fixes that. Every past day still counts forever, it just counts less and less the further back it is. You have already met this idea. It is the EMA, exponential moving average, sitting on every trading chart you have opened.
The formula, one Greek letter at a time
Simple Exponential Smoothing keeps a running number called the level. Each new day updates it:
new level = alpha × today's actual demand + (1 − alpha) × yesterday's level
tomorrow's forecast = new level
Alpha is a single number between 0 and 1 that decides how much weight lands on today versus everything that came before. Let's watch it work by hand before touching real data. Say alpha is 0.5, and yesterday's level, your running estimate of "where demand normally sits," was 100 kg.
| Day | Actual demand | Yesterday's level | New level = alpha × actual + (1 − alpha) × old level |
|---|---|---|---|
| 1 | 110 | 100 | 0.5×110 + 0.5×100 = 105 |
| 2 | 90 | 105 | 0.5×90 + 0.5×105 = 97.5 |
| 3 | 108 | 97.5 | 0.5×108 + 0.5×97.5 = 102.75 |
Each day's forecast for tomorrow is simply today's new level. Notice the level never jumps all the way to the newest actual, it only moves halfway there each time, since alpha is 0.5. With alpha at 0.9 it would leap most of the way to the newest number every day. With alpha at 0.1 it would barely budge, holding close to its old estimate for a long time. That is the entire mechanism, repeated for as many days as you have data. In the trading world, an EMA's length in days converts to alpha by alpha = 2 ÷ (length + 1), so a short, twitchy EMA is a high alpha, and a long, calm EMA is a low alpha. Same knob, same trade-off you already understand: fast and jumpy against slow and stable.
Slide alpha all the way to 1.00 and check the four numbers against module 2's naive scoreboard: 29.7, 36.0, 1.23, minus 0.2. Identical. At alpha equals 1 the level forgets everything except yesterday, which is exactly what naive does. SES is not a different family of ideas, naive and the moving average are both special cases sitting inside it.
Holt and Holt-Winters: adding trend and season
Plain SES has exactly one blind spot: it only ever tracks "roughly where demand sits," a single number. If demand is steadily climbing by about 5 kg a day, SES's level will always be a step behind, chasing a moving target, because it was never given anywhere to store "how fast is this thing moving." Holt's method fixes that by keeping a second running number just for that, the trend, updated by its own smoothing weight so it too can be fast-moving or slow-moving. Holt-Winters, sometimes called triple exponential smoothing, adds a third running number, the season, for a shape that repeats on a fixed cycle, like cafe demand being reliably higher every Saturday. Three separate jobs, three separate running numbers, three separate smoothing weights so each can forget the past at its own pace:
| Component | Formula | In plain words |
|---|---|---|
| Level | L(t) = α(Y(t) − S(t−p)) + (1−α)(L(t−1)+T(t−1)) | Where the series sits right now, seasonality stripped out |
| Trend | T(t) = β(L(t) − L(t−1)) + (1−β)T(t−1) | How fast the level itself is drifting up or down |
| Season | S(t) = γ(Y(t) − L(t)) + (1−γ)S(t−p) | The repeating bump or dip for this point in the cycle |
Alpha, beta, gamma. Each is its own dial between 0 and 1, each controls how fast that one component forgets the past. p is the length of one full seasonal cycle, 7 for a weekly pattern.
Additive or multiplicative
A seasonal bump can behave two ways. Additive means the bump stays a fixed size no matter how big the series gets, level plus trend plus season. Multiplicative means the bump scales with the level, (level plus trend) times season. Cafe weekends are the real test: across the actual Cafe dataset, weekday demand averages 975 units and weekend demand averages 1,085, a bump of about 110 units. If overall traffic doubled some year, would that weekend bump still be a flat 110 units, or would it roughly double too? For a cafe, more customers behaving in the same weekend pattern reads as proportional, which points toward multiplicative. Fit both to the real data and see which one the numbers actually prefer.
| Model, fit on real cafe data | AIC | MAPE % |
|---|---|---|
| Additive | 6085.4 | 4.72 |
| Multiplicative | 6086.3 | 4.73 |
Honest result: on this real dataset the two are nearly tied, AIC differs by under 1 point, MAPE by one hundredth of a percent. No real conflict here. Both criteria quietly agree: go additive, chiefly because it is the simpler model and the data gives no reason to pay for the extra complexity.
When AIC and MAPE actually disagree
Before the names, the actual problem these numbers solve. Give a model more knobs to turn, more terms, more smoothing weights, more everything, and it will almost always fit the training data it already saw better, that is just arithmetic, more flexibility can always bend closer to points you already know. But fitting the past well and forecasting the future well are not the same skill: a model with enough knobs can bend itself around pure noise, memorising old wiggles that will never repeat, which makes it worse at the one job that actually matters, forecasting days it has not seen yet. So you need a way to reward a model for fitting well while also making it pay a toll for every extra knob it needed to get there.
That toll is exactly what AIC and BIC are. Full names: AIC is Akaike Information Criterion, BIC is Bayesian Information Criterion, also written SBC for Schwarz Bayesian Criterion. Both start from how well a model fits, then subtract a penalty sized to how many parameters it used, lower is better on both. BIC's penalty grows more harshly with the amount of data you have, so it leans toward even simpler models than AIC does once your dataset is large. MAPE, mean absolute percentage error, is a completely different kind of check: it never looks at the training fit at all, it only scores how well the model predicted days it never trained on, the actual, direct test of forecasting skill. AIC and BIC usually agree with MAPE, since a model that overfits training noise both racks up complexity penalty and forecasts badly on new data. When they do not agree, here is the professor's own worked conflict and his three-level way out of it.
| Metric | Additive | Multiplicative | Difference |
|---|---|---|---|
| AIC | 1120 | 1130 | 10, additive wins |
| MAPE | 9.2% | 8.7% | 0.5%, multiplicative wins |
Level 1: what is the use case
Forecasting the future, prioritise MAPE. Explaining the past to stakeholders, prioritise AIC. Running unattended at scale, prefer the simpler model unless the gain is large.
Level 2: how big is the gap
An AIC gap past 10 points is substantial. A MAPE gap under 1 percent is not worth switching over. Here the AIC gap is exactly 10 and the MAPE gap is only 0.5, so level 2 alone does not settle it.
Level 3: explainability
In operations the simpler model wins unless the more complex one earns its keep. Choose additive here, since the multiplicative gain is small and the AIC penalty against it is not.
As analysts your job was never to chase the lowest number on a screen. It is to balance fit, accuracy, simplicity, and whether the business can actually act on what you built.
On the SES bench above, why does a very low alpha carry a large positive bias for the first part of the test window that shrinks toward zero as alpha rises, when the naive-equivalent alpha of 1.0 has almost no bias at all? And separately: a colleague reports AIC favours model X by 3 points, MAPE favours model Y by 2 percent, and the forecast will run unattended feeding a nightly production plan. Which model do you ship?
Low alpha means the level barely moves from wherever it was seeded, the last training day's value. If that seed happened to sit a little above the test month's true average, and it did here, a slow-moving level keeps leaning on that high value for weeks before it catches up, producing a sustained positive bias. A high alpha updates almost every day to the newest actual, so it never drifts far from wherever demand actually is, at the cost of a noisier line. Smoothness and staying unbiased are in tension, not the same goal.
Ship model Y. The AIC gap of 3 is well under the substantial threshold of 10, so level 2 says it does not matter much. The forecast running unattended in production is exactly the use case where MAPE, real out-of-sample accuracy, should decide, per level 1. And model Y already has the accuracy edge, so there is no level 3 trade-off to even weigh.
ARIMA, ARIMAX, SARIMA
Every morning at BrewVista Cafe, Ravi has to decide how much to prepare before he knows what today's actual demand will be. Prepare too little and he loses sales and disappoints customers. Prepare too much and it is waste, straight cost. This is not a statistics exercise dressed up as a story, it is the actual operating decision the rest of this module is built to support.
Four forecasting assistants, each adding one source of information
Full name first, since it earns one: ARIMA is AutoRegressive Integrated Moving Average. Three ingredients, each with its own letter and its own job. AR, autoregressive, uses the variable's own past values. I, integrated, means differencing the series to strip out trend. MA, moving average, uses past forecast errors, not the moving-average-of-values from module 2, a genuinely different idea wearing the same two words. From there, three more assistants each add one new ingredient on top:
ARIMA
Learns purely from Demand's own past values and past errors. No outside information.
ARIMAX
Adds outside variables: temperature, promotion, weekend, price. The X stands for exogenous.
SARIMA
Adds a seasonal memory: what happened on this same weekday last week, and the week before.
SARIMAX
History, seasonality, and outside drivers, all at once.
Fit all four on the real cafe dataset, holding out the last 30 days, and the improvement is not theoretical:
| Model | MAE units | RMSE units | MAPE % |
|---|---|---|---|
| ARIMA, history only | 71.1 | 82.8 | 7.17 |
| ARIMAX, + drivers | 51.0 | 59.9 | 5.05 |
| SARIMA, + seasonality | 47.2 | 61.5 | 4.7 |
| SARIMAX, everything | 34.6 | 42.7 | 3.43 |
Each assistant genuinely earns its place here: error keeps falling every time one more source of real information gets added.
Before any of this: is the series stationary
Stationary means a series' statistical personality does not change over time: constant mean, constant spread, and whatever is left over after removing structure looks like plain noise, not a lingering pattern. Trend and seasonality are the two usual reasons a series fails this test. You already have the trading equivalent: this is the same question as asking whether a spread is mean-reverting before you trust a pairs trade on it.
The test itself is the ADF test, Augmented Dickey-Fuller, named after the two statisticians who built it. It hands back a single number called a p-value, and that number deserves a proper explanation before you ever read one, since almost every confusing moment with statistics traces back to this exact idea.
A p-value answers one specific question: assume, just for the sake of argument, that the series is secretly not stationary. Given that assumption, how surprising would it be to see data that looks this calm and settled purely by chance? A small p-value, under 0.05 by the usual convention, means the answer is "very surprising, this would be a real coincidence if it truly weren't stationary," so you conclude it probably is stationary after all. A large p-value means the opposite: nothing here rules out the "not stationary" assumption, so you cannot confidently call it stationary. The number 0.05 itself is just a common convention, roughly "willing to be wrong about this call one time in twenty," not a law of nature.
Differencing, the usual fix when a series fails this test, is nothing more than subtracting each value from the one right before it. A tiny worked example, five made-up days of a slowly climbing series:
| Day | Raw value | Differenced, today minus yesterday |
|---|---|---|
| 1 | 100 | - |
| 2 | 105 | 105 − 100 = 5 |
| 3 | 103 | 103 − 105 = -2 |
| 4 | 110 | 110 − 103 = 7 |
| 5 | 108 | 108 − 110 = -2 |
The raw column is climbing overall, 100 up to 108. The differenced column, 5, -2, 7, -2, has no climb left in it at all, it just wobbles around a flat average near zero. That is exactly what differencing is for: it throws away the level and keeps only the day-to-day change, which strips out a trend that would otherwise make the mean drift over time and fail the stationarity test.
A p-value of 0.0641 sits right on top of the usual 0.05 line, genuinely too close to call by eye, which is exactly why this dataset carries the name it does. One round of ordinary differencing settles the question completely, the statistic collapses and the p-value effectively hits zero. When in doubt, difference and re-test, never eyeball a line chart and declare it stationary.
Reading a lagged value
A lag is simply an earlier reading of the same series, lined up next to today's. Lag 1 is yesterday, lag 2 is the day before that. Small worked example with daily temperature:
| Day | Temperature | Lag 1 | Lag 2 |
|---|---|---|---|
| 1 | 20 | - | - |
| 2 | 22 | 20 | - |
| 3 | 21 | 22 | 20 |
| 4 | 23 | 21 | 22 |
Correlation itself, before the two long names: it is a single number between -1 and +1 that measures how much two lists of numbers move together. Close to +1 means when one is high the other tends to be high too. Close to -1 means the opposite, one high tends to pair with the other low. Close to 0 means no real relationship either way. ACF, Autocorrelation Function, and PACF, Partial Autocorrelation Function, both apply that exact same idea to a series and its own past self, "auto" meaning self. ACF asks the raw question, does today correlate with 7 days ago at all, no matter how it got there. PACF asks a stricter version of the same question, after already accounting for every shorter lag in between, isolating what lag 7 specifically adds on its own, once lags 1 through 6 have already had their say. The classic Box-Jenkins rule of thumb, named for the two statisticians who popularised this method: PACF points to p, the AR order, ACF points to q, the MA order, each read from wherever the plot cuts off sharply to near zero.
The trap: not every spike is a new order
Difference a series with weekly structure once, ordinary differencing, d equals 1, and its ACF plot often still shows large spikes at lag 7, 14, 21, 28, 35, 42. The tempting, wrong conclusion is to read six spikes as q equals 6. The correct read: ordinary differencing only removed trend-like movement, weekly dependence is still sitting there untouched. Spikes repeating at multiples of 7 are the signature of seasonality, not six extra moving-average terms. The fix is seasonal differencing, not a bigger q.
| Compares today to | Notation | |
|---|---|---|
| Ordinary differencing | Yesterday | d = 1 |
| Seasonal differencing | Same weekday last week | D = 1, s = 7 |
| Combined | Both at once | d = 1, D = 1 |
SARIMA writes this as SARIMA(p,d,q)(P,D,Q)s. The letters in parentheses without a subscript are the familiar nearby-lag terms. The letters with the subscript s are the same three ideas applied at seasonal distance, s equals 7 for a weekly pattern, so seasonal lag 1 in that notation actually means calendar lag 7.
Rebuilding your professor's own order search
Same Ambiguous Case data, same five candidate orders, checked by fitting each on the training data and scoring it on 30 held-out days:
| Order (p,d,q) | AIC | BIC | RMSE |
|---|
This independent run lands within rounding error of your professor's own notebook output on the same file, order (1,1,1) wins on every column at once, which does not happen often enough to take for granted.
The one rule that actually matters for ARIMAX
If X is used in the model, future X must be supplied to forecast with it. Weekend and price are usually knowable in advance. A promotion is planned by the business itself. Temperature needs an actual weather forecast. This is not a footnote, it decides whether ARIMAX can even run for a given horizon.
Choosing between the four, in order
Never choose by AIC alone. The professor's actual sequence, in priority order:
| Step | Check |
|---|---|
| 1. Residuals | Ljung-Box test, named for Greta Ljung and George Box, checks whether what is left over is just noise. p above 0.05 is acceptable, at or below it means real structure remains unexplained. |
| 2. Holdout accuracy | Compare RMSE, MAE, MAPE on the test period, never on data the model already trained on. |
| 3. AIC / BIC | Supporting evidence only, within the same sample. |
| 4. Operational usability | Can you reliably get future X values when it actually has to run next month. |
| 5. Parsimony | If it is still close after all of that, take the simpler model. |
A lower AIC does not guarantee the better forecast on data the model has never seen. Choose ARIMAX when drivers matter and their future values are actually available. Choose SARIMA when seasonality is strong and those drivers are not available or not needed. Choose SARIMAX when both are true at once. Reject and rethink whenever residuals still show a pattern, or the model needs an input you cannot reliably supply in production.
A colleague differences a weekly-patterned series once, d = 1, sees ACF spikes at lag 7, 14, 21 and 28, and proposes an MA(4) model. What is wrong with that read, and what should they try instead? Separately: SARIMAX comes back with the lowest AIC of all four assistants on your test set, but it needs tomorrow's temperature, promotion plan and price index to produce a forecast. Under what condition would you still reject it in favour of SARIMA?
Spikes at multiples of 7 after ordinary differencing are the signature of leftover weekly seasonality, not four separate moving-average relationships. Counting each spike as a new q term is the exact trap this module warned about. The fix is seasonal differencing, D = 1 with s = 7, not a bigger MA order.
Reject SARIMAX for SARIMA whenever the future exogenous values cannot be trusted in production, no reliable temperature forecast, no confirmed promotion calendar, no locked price plan. This module's own real numbers prove the cost of getting that wrong: feeding ARIMAX frozen, incorrect future exogenous values pushed its error to worse than plain ARIMA, which uses no outside information at all. A model that scores best on paper but cannot be fed correctly in production is not the better model, it is a liability with a good AIC.
Forecast to decision
A forecast that never turns into a decision is trivia. This module is the payoff for everything so far: how much to actually produce or stock, given a forecast you now know how to build and how to judge.
Why inventory exists at all
A typical firm holds roughly 30 percent of its current assets, and as much as 90 percent of its working capital, in inventory. That is not an accident, inventory exists on purpose: to meet demand it cannot serve instantly, to smooth production so a factory is not restarting every day, to decouple one stage of operations from the next, and to protect against running out. Every one of those reasons costs money to maintain, which is the entire tension this module is about.
Sorting inventory so effort goes where it matters
Two classification systems actually get used downstream, in module 6's automation. ABC sorts items by annual value, not by how many there are: A items are typically only 10 to 20 percent of item count but 60 to 70 percent of annual value, and earn the tightest control. C items are 50 to 60 percent of item count but only 10 to 15 percent of value, and do not deserve the same attention. FSN instead sorts by how often an item actually moves: Fast-moving needs constant replenishment, Slow-moving may need phasing out, Non-moving has not sold at all and may need writing off. A single item carries both tags at once, an A-category item can still be Non-moving, which is usually a red flag worth a second look.
How much to order: EOQ
EOQ, Economic Order Quantity, answers a different question from "how much safety stock." It asks: given that ordering itself costs money, every purchase order has an administrative cost, and holding stock costs money too, capital tied up, storage, insurance, what order size minimises the two added together. Order too often and ordering cost piles up. Order too rarely in huge batches and holding cost piles up. The formula itself, D for annual demand in units, S for the cost of placing one order, H for the cost of holding one unit for a year:
EOQ = √(2 × D × S ÷ H)
A tiny worked number so the formula is not just symbols: a shop sells 1,000 units a year, each order costs ₹50 to place, and holding one unit for a year costs ₹2. EOQ = √(2 × 1,000 × 50 ÷ 2) = √50,000 ≈ 224 units per order. Order much less than 224 at a time and you are paying ₹50 too often. Order much more and you are paying ₹2-per-unit holding cost on stock that just sits there. The total cost curve traced across every possible order size is U-shaped, and its minimum sits exactly where the ordering-cost curve and the holding-cost curve cross, which is precisely what that square root is computing. There is a genuine sweet spot, not "more is better" and not "less is better."
Safety stock and the reorder point
Reorder point is the stock level that triggers a new order. Under demand uncertainty, it has two pieces, and one of them, Z, needs unpacking before the formula means anything:
Safety stock = Z × σ(daily demand) × √(lead time)
Reorder point = (average daily demand × lead time) + Safety stock
Standard deviation, the σ in that formula, measures how spread out daily demand usually is around its own average. A small standard deviation means most days land close to average. A large one means days swing far from it in both directions. If you plot many days of demand as a histogram, it typically piles up into a bell shape, most days near the average, fewer and fewer the further you go from it in either direction, and that bell shape is called the normal distribution. A Z-score is just a way of saying "how many standard deviations above the average," and because the bell shape is so consistent, a given Z-score always covers the same fixed share of days. Z of 1.65 covers 95 percent of the area under that bell: in other words, if demand really does follow this bell shape, only the worst 5 percent of days will exceed 1.65 standard deviations above average. Carry enough safety stock to cover demand up to that line, and you will run out only on one of those rare worst days. That is the entire meaning behind "Z of 1.65 means 95 percent service." Push Z higher and you are drawing the line further out into the bell's thin tail, stockout risk falls, but the safety stock, and its holding cost, needed to reach that far out both climb.
Real case data: average daily demand 100 units, standard deviation 20, lead time 5 days, holding cost ₹20 per unit per year, order quantity 1,000 units fixed by the supplier, profit lost per unit stocked out about ₹150.
Click through all four. The cheapest option here is not 99 percent service, it is 98. Past that point, the safety stock you are carrying every single day costs more than the rare stockout it prevents. Your professor's own group question makes the harder point: a model built for 95 percent service still produced stockouts in over 20 percent of simulated cases. Real daily demand is rarely a clean, symmetric normal curve, and a formula this simple assumes it is.
Back to Mahakoshal: what a buffer decision is actually worth
Module 2 ranked three forecasts by accuracy alone. Here is what that ranking is actually worth in rupees, running the real inventory-carryover simulation your course code uses: forecast plus a buffer, clipped to the 2,400 kilogram daily capacity, carried forward as opening inventory the next day, shortage costed at ₹5 per kilogram, holding at ₹1 per kilogram per day.
| Forecast | Buffer kg | Shortage kg | Holding kg-days | Total cost ₹ |
|---|
Two things worth noticing. At buffer plus 100 every model ties at ₹1,320: the buffer pushes production against the 2,400 kilogram ceiling almost every day regardless of which forecast fed it, so the capacity constraint, not forecast accuracy, becomes the thing actually deciding the outcome. At buffer zero the three forecasts finally separate, and MA_7 wins clearly, the same model module 2 crowned on MAE alone. At buffer minus 100, deliberately under-producing in a plant that is already barely sufficient, costs explode past ₹14,000 because inventory never gets the chance to build a cushion.
At buffer plus 100, why does the choice of forecasting model stop mattering, when modules 2 and 3 spent so much effort separating naive from MA_7 from SES on accuracy alone? What does that tell you about when forecast accuracy is actually the binding constraint on a decision, and when something else is?
Average demand here sits at roughly 2,402 kilograms, almost exactly at the 2,400 kilogram capacity. Add a positive buffer to any of the three forecasts and forecast-plus-buffer nearly always exceeds capacity, so production gets clipped to 2,400 regardless of which model produced the underlying number. Once the ceiling is doing all the work, a more accurate forecast underneath it cannot buy you anything more, there is no room left to act on the extra accuracy.
The general lesson: forecast accuracy only pays off when the decision has room to actually use it, when you are not already pinned against a hard constraint like plant capacity. Improving a forecast that feeds a capacity-constrained decision is effort spent in the wrong place, the real lever there is capacity or the buffer policy itself, not the model.
Process automation
High SKU count plus demand that will not sit still turns manual tracking into a losing game: restocking runs late, reorder triggers get missed, and the business swings between piling up stock and running out. The fix is not a bigger ERP system, it is a small, automated, supplier-aware layer on top of what already exists. If you were building this as a feature in one of your own apps, this is exactly the spec you would write.
The stack, three cheap tools doing one job each
Google Sheets + Apps Script
The trigger and the messenger. Reads the live stock sheet, decides who needs an alert, sends it.
Python notebook
The classifier. Runs ABC and FSN, and the network-based transport decisions from module 7.
Gephi
Pure visualisation of the supply chain as a graph, nothing computed here, just seen clearly.
The actual workflow
Read the inventory sheet. Group items by supplier, since a restock call happens supplier by supplier, not item by item. Flag any item where Current Stock has dropped below its Reorder Level. For each flagged item, compare its Expected Stockout Date against its Projected Receipt Date, if the shipment is due to land after the item is already expected to run out, that is real danger, not just low stock. Compose one HTML table per supplier and email it out automatically, stamped with the run time. All of it lives in one function, sendInventoryAlerts, plus a small formatDate helper, both running inside the spreadsheet itself.
By supplier
| Supplier | Items below reorder | In real danger | Danger rate |
|---|
Five items that would be in this morning's email right now
| Item | Supplier | ABC | FSN | Stock | Reorder at | Stockout | Receipt |
|---|
Supplier Z has fewer items below reorder than Supplier Y. Which supplier's phone call should actually go out first this morning, and what number from the table above justifies it?
Supplier Z, despite having the fewest items below reorder of the three. Look at the danger rate column, not the raw count: a much higher share of Supplier Z's understocked items are genuinely going to stock out before the replacement shipment lands. Raw count tells you who has the most open items. Danger rate tells you whose problem is actually urgent. An automated queue sorted only by count would call the wrong supplier first.
Network design and optimization
NavDisha Consumer Products ships from 2 plants through 3 warehouses to 6 retail markets. Distribution cost has crept up and several markets depend on a single warehouse connection. Management wants two things: understand the structure of the network before touching it, then decide exactly how much product should move over which route, at the lowest total cost.
One deliberate trap sits inside the actual route data: alongside real cost there is a "proximity weight," equal to 1 divided by distance, meant only for visualising closeness. It is not a cost to minimise. Before trusting any result, check which column a method actually used.
Step one: diagnose the structure, before any optimizing
Degree centrality counts how many direct connections a node has, out of every connection it could possibly have. Betweenness centrality counts how often a node sits on the shortest path between two other nodes, which is what actually flags bridging points and single points of failure. A node can be structurally central while carrying little flow, and a peripheral node can still carry heavy demand, the two questions are not the same. A tiny made-up network makes both concrete before the real one: one hub, H, connected to four spokes, A, B, C and D, with no direct connections between any two spokes at all, everything has to pass through H.
Degree centrality: H connects to all 4 other nodes, the most it possibly could, so its degree centrality is as high as it gets. Each spoke connects to only 1 node, H itself, so each spoke's degree centrality is low. Betweenness centrality: to get from A to B, the only path runs through H, so H sits on that shortest path. The same is true for every other pair of spokes, A to C, A to D, B to C, B to D, C to D, H sits on all of them. A spoke, meanwhile, never sits on the shortest path between two other nodes, since nothing routes through a spoke. So H scores the maximum possible on both measures at once, and every spoke scores zero on betweenness even though each spoke still has real demand of its own. Hold that shape in mind, hub high on both, spokes high on neither, it is exactly the pattern the real network below produces.
| Node | Type | Degree centrality | Betweenness centrality |
|---|
Indore tops both columns at once. It is a genuinely useful consolidation point and a disruption risk in the same breath: Nashik and Ujjain have no other warehouse to fall back on if Indore ever goes down.
Community detection is structure, not a distribution plan
Running community detection on this exact network finds three tight clusters. Two of them look like a complete plant-warehouse-retailer chain. The third does not:
| Community | Members | Contains a plant? |
|---|---|---|
| 1 | Ahmedabad (plant), Surat (warehouse), Vadodara, Rajkot (retailers) | Yes |
| 2 | Indore (warehouse), Bhopal, Nashik, Ujjain (retailers) | No plant |
| 3 | Pune (plant), Nagpur (warehouse), Mumbai (retailer) | Yes |
Community 2 has no plant of its own at all. A rule that forces every retailer to be served only from plants and warehouses inside its own community would leave Bhopal, Nashik and Ujjain completely unservable, a real structural cluster is not the same thing as a self-sufficient distribution territory.
Step two: the actual linear program
A linear program is just three plain-English pieces dressed up in maths notation. The decision variables are the actual numbers you are free to choose, here, how much to ship on each individual route. The objective is the one number you are trying to make as small (or as large) as possible, here, total cost. The constraints are the rules your choice is not allowed to break, plant capacity, retailer demand, warehouse balance. Solving the program means searching every combination of shipment quantities that obeys every constraint, and keeping the one with the lowest total cost, which is exactly the same kind of search Excel's Solver or a portfolio optimiser runs, just with shipping routes standing in for stocks. For this network specifically: the decision variables are how much flows on each plant-to-warehouse and warehouse-to-retailer route. The objective is total cost, quantity times the real unit cost, never the proximity weight. The three constraints: no plant ships more than its capacity, every retailer's demand is met exactly, and every warehouse's inbound flow equals its outbound flow, nothing accumulates inside a warehouse in this model.
| Route | Flow, units |
|---|
When does a shortcut heuristic actually match the optimum
Running the simple version of the greedy rule, serve each retailer through whatever feasible route is cheapest, one retailer at a time, on this exact base case produces ₹39,000 serving all 2,000 units, identical to the linear program. That is not a coincidence to expect everywhere, it happens here because each retailer realistically has very few routing choices and no plant capacity is tightly binding in the base case. Push capacity harder, as the stress case above does, and a greedy rule's blind spots start to show: it commits early to a cheap-looking route and has no way to revisit that choice once a later retailer turns out to need it more.
Toggle to the capacity-shock case above. Which plant-to-warehouse route changed, and why that one specifically rather than some other route in the network?
Pune, previously shipping only to Nagpur and sitting on 1,000 units of unused capacity in the base case, starts shipping 500 units to Indore as well. Ahmedabad was Indore's cheaper supplier before the cut, at a unit cost of 10 against Pune's 14, but Ahmedabad alone can no longer cover both Indore and Surat once its capacity drops to 1,000. The solver reaches for the next-cheapest feasible source, which is Pune, and it can do that specifically because Pune had slack capacity sitting idle in the base case. A plant running at the edge of its own capacity would have had nowhere to absorb the shock from.
What is deliberately missing
Every module above is reverse engineered from files that actually exist in the course folder: the syllabus, four PPT decks, the lecture code, and the real datasets. Sessions 10, and 13 through 20, contracts, warehousing detail beyond module 5, the integer and binary optimisation sessions, the two integrated cases, and the Gen AI sessions, have no source files yet, so nothing here invents them. A Session 11-12 folder does exist with a scaled-up version of module 7's network, 20 plants, 50 warehouses, 500 retailers, 2,800 routes, sitting ready as a stretch exercise once module 7 feels solid.
Built alongside the chat session it came from. Every number on this page is computed directly from the real course files, not invented for teaching: Session_01_Large_Yarn_Demand_Dataset.xlsx (10,008 hourly readings, 417 days), the Ambiguous ARIMA Case, the Cafe ARIMAX/SARIMA dataset, Final Inventory Alerts, and the NavDisha transportation caselet. Course: Operations Analytics 1, Prof. Akshay Khanzode, IIM Visakhapatnam.