On the Lean Toolkit
ANNOTATED
Eight tools born on factory floors that get taught as a certification and used as a poster on a wall. The printed text is the clean definition each one earns in the classroom. The margins — and the desk notes — are what happens when you take the same tools out of the factory and run them in a flooring business, an agency, a SaaS rollout: anywhere people, deadlines, and pride are in the room at once.
Muda — the Seven Wastes
Muda is the foundation the rest of the toolkit stands on, because every other Lean tool is, at bottom, a specific technique for removing a specific kind of waste. The discipline begins by teaching the practitioner to see it. The seven wastes — often taught with a mnemonic — are the categories into which almost all value-destroying activity falls, and learning to name them is the first step toward removing them.
The most important of the seven is generally held to be overproduction — making more than is needed, sooner than it is needed — because it is the waste that generates the others. Overproduction creates inventory that must be stored (waste), moved (waste), and inspected, and it hides defects inside a batch that will not be discovered until much later. To attack overproduction is therefore to attack the root that feeds the rest.
Crucially, Lean distinguishes muda from the two related concepts of mura (unevenness) and muri (overburden). Waste is often the visible symptom; unevenness in flow and overburden on people or machines are frequently its underlying causes. Removing a visible waste without addressing the unevenness that produced it guarantees the waste returns.
The seven wastes were codified inside Toyota by Taiichi Ohno, who reportedly could stand on a factory floor and see waste others walked past. The entire Toyota Production System — the reason a post-war Japanese carmaker overtook giants many times its size — rests on relentlessly hunting muda. The lesson travels: Toyota didn't win on a better engine. It won on a better process for removing waste, repeated for decades.
Muda in a home-services business looks nothing like a factory and costs just as much: a second truck roll because the first visit missed a measurement, materials handled twice between supplier and site, a crew waiting on an approval that sat in someone's inbox. And the eighth waste is real here too — the day an installer's process suggestion gets ignored is the day the ideas stop. — A.P.
Eliminate waste. Eliminate the unevenness that keeps manufacturing the waste. Symptoms regrow; causes don't.
Kaizen — Continuous Improvement
Kaizen holds that the sum of many small improvements, made continuously by the people closest to the work, outperforms the occasional grand initiative imposed from above. Its power is compounding: a process improved by a fraction of a percent every week is transformed within a year, and — more importantly — the capability to improve becomes embedded in the organisation rather than dependent on outside intervention.
The mechanism is typically a structured cycle — plan, do, check, act — repeated indefinitely. A small change is proposed, tested on a limited scale, measured against the prior state, and either adopted and standardised or discarded. The standardisation step is essential: an improvement that is not written into the standard way of working is an improvement that will quietly decay back to the old method.
Kaizen is frequently contrasted with kaikaku — radical, discontinuous change. The two are not opposites but complements: kaikaku resets the baseline with a large step, and kaizen relentlessly improves from the new baseline. An organisation that only knows how to do large transformations will stagnate between them; one that only does kaizen may never make a leap the situation demands.
Jeff Bezos built Amazon on a refusal to ever declare the process finished — his famous "it's always Day 1" is kaizen written as company religion, the belief that "Day 2 is stasis, followed by death." The compounding flywheel — lower prices bring customers, customers bring sellers, scale lowers costs again — is exactly kaizen's logic: small improvements that feed each other and compound. Note: the flywheel only spins because every turn is measured and standardised, not because anyone hustled harder.
Kaizen for me is a Friday habit, not a factory programme: one small tweak to the quoting template, the scheduling rules, or the client-onboarding email every week — tested, and if it works, written into the standard doc that day. In the agency we run the same loop as a ten-minute retro after every campaign. Small, boring, relentless. It compounds. — A.P.
Standardise every improvement, or it decays. And close the loop on every suggestion, or the suggestions stop.
- "Waste" is too soft a word. Muda is time, effort, and cash stolen from someone downstream. Name it that.
- The eighth waste is the worst. Unused human potential — a floor that stopped suggesting fixes — dwarfs the material seven.
- Kaizen trusts the person doing the job. They're the genuine expert on where it breaks. That's the whole idea.
- Kaizen dies on ignored ideas, not bad technique. Close the loop fast, or the suggestions stop forever.
Poka-Yoke — Mistake-Proofing
Poka-yoke rests on a realistic view of human attention: people, however skilled and well-intentioned, will occasionally err, and no amount of training or exhortation reduces that rate to zero. Rather than demanding perfect vigilance, poka-yoke redesigns the task so the error cannot occur — a connector that only fits the correct way, a form that will not submit until a required field is filled, a step that physically cannot proceed until the prior one is complete.
The discipline distinguishes between prevention and detection. A prevention poka-yoke makes the error impossible; a detection poka-yoke makes it immediately visible so it can be corrected before it travels downstream. Prevention is always preferable, because a detected error has still consumed effort, but detection is a valuable fallback where prevention is impractical.
Well-designed poka-yoke is typically simple and inexpensive — a guide pin, a checklist that gates the next step, a warning light. The elegance of the concept is that it moves the cost of quality from ongoing inspection, which is a permanent tax, to a one-time design change, which pays out indefinitely.
The best mistake-proofing is invisible because it simply works. A car won't start unless it's in park; an ATM returns your card before dispensing cash, so you can't walk away without it; a SIM card only fits one way; a surgical checklist won't let the team proceed until the count is confirmed. None of these rely on anyone being careful. That's the tell: the companies obsessed with quality didn't hire more vigilant humans — they designed workflows where the wrong action simply won't complete.
My favourite poka-yokes have no machines anywhere near them: a quote form that won't submit without site photos and measurements, a job that can't be scheduled until the deposit clears, a CRM field that's required before a lead can move stages. Nobody has to remember anything — the workflow simply refuses to proceed wrong. — A.P.
Train people not to make the error. Design the process so the error can't be made. Training expires; a guide pin doesn't.
Value Stream Mapping
Value stream mapping forces a whole-system view. Rather than optimising individual steps in isolation — the local optimisation warned against in Chapter I — it lays out the complete journey from raw input to delivered value, annotating each step with the time it takes and the time spent waiting between steps. The map makes visible what intuition consistently gets wrong: where the time actually goes.
The technique separates the current-state map from the future-state map. The current state records how the process actually behaves today, including every delay, handback, and queue. The future state is the redesigned flow the improvement effort aims at. The gap between the two becomes the improvement plan, and the map turns a vague sense that "things are slow" into a specific, located set of problems.
Because the map includes the flow of information as well as material, it exposes a class of problem that step-level thinking misses entirely: the delays caused by how work is scheduled, approved, and handed off. Frequently the largest opportunities are not in the physical work at all but in the information flow that governs when the work is allowed to proceed.
When a pizza chain promises "thirty minutes," it has implicitly value-stream-mapped its own operation: order capture, make-line, oven, cut-and-box, out-the-door. The promise is only keepable because someone mapped where the minutes actually go and attacked the queues — not by telling cooks to move faster, but by redesigning the make-line so a pizza never waits. The general truth: any business that competes on speed has, knowingly or not, mapped its stream and killed the waiting between steps.
Map enquiry-to-installed-floor and the result is humbling: the labour is days, the lead time is weeks — the gap is waiting on approvals, supplier ETAs, and unanswered quotes. Same in agency work: the shoot takes a day, the project takes a month, and the difference is queues. I attack the waiting now, and mostly leave the work alone. — A.P.
Attack the waiting, not the work. The queue between steps is where the lead time lives.
- Ask "how did the process allow it," not "who did it." Poka-yoke is that question turned into physical design.
- "Be more careful" is the absence of a fix. Design the error out or expect it back on the next bad day.
- Value-add is a tiny fraction of lead time. The work takes minutes; the waiting takes days. Attack the waiting.
- Map the real process, not the SOP. The gap between official and actual is where the big problem hides.
Kanban — Pull, Not Push
Kanban replaces the push logic of "produce as much as you can" with the pull logic of "produce only what the next step is ready to consume." A signal — historically a physical card, now often a card on a board — travels back up the line to authorise the next unit of work. Nothing is made until something downstream has pulled, which directly attacks the overproduction identified in Article 1 as the root waste.
The visual nature of the system is central. Because the state of all work is displayed, bottlenecks announce themselves: work visibly piles up in front of the constrained step. This makes Kanban a diagnostic instrument as much as a scheduling one — it does not merely control flow, it reveals exactly where flow is failing, in real time, to anyone who glances at the board.
Kanban's discipline is limiting work in progress. By setting an explicit ceiling on how many items may occupy any stage at once, the system forces the completion of existing work before new work is begun. This exposes bottlenecks that an unlimited push system would simply bury under an ever-growing pile of partially finished inventory.
Kanban's origin is oddly domestic: Ohno took the idea from American supermarkets, where a shelf is only refilled when a customer removes an item — the shelf "pulls" stock rather than the back room "pushing" it. Decades later the same logic runs modern software: teams at companies from Spotify to countless startups manage work on kanban boards with strict WIP limits. What survived the jump from grocery aisle to code sprint wasn't the cards — it was "don't start new work until you've finished what you pulled."
I run WIP limits on renovation jobs and on content projects alike: only so many active at once, full stop. The counter-intuitive rule holds outside every factory I've seen — the month we capped active jobs, completions went up, callbacks went down, and the schedule finally told the truth. — A.P.
Visualise the work on a board. Limit the work in progress. The board is the thermometer; the WIP limit is the medicine.
SMED — the Changeover Tax
SMED addresses a cost that is easy to overlook because it produces nothing: the time a process spends stopped while it is reconfigured from one job to the next. Long changeover times have a hidden strategic consequence — they push a business toward large batches, because if changing over is expensive, the instinct is to change over as rarely as possible and run long. Large batches then reintroduce every waste of overproduction.
The core insight of SMED is the distinction between internal and external setup. Internal setup can only be done while the process is stopped; external setup can be done while the process is still running the previous job. The single largest gain usually comes not from doing the setup faster, but from converting internal steps into external ones — preparing everything possible in advance so the actual stopped time shrinks dramatically.
SMED thus enables the flexibility that pull systems and small batches require. It is the tool that makes responsiveness affordable: when changeover is cheap, a business can produce exactly what is needed in the quantity needed, rather than committing to long runs it must then store, move, and eventually discount.
An F1 pit stop is SMED taken to its limit: a tyre change that once took a minute now takes under three seconds, achieved almost entirely by moving work external — every tool staged, every crew member positioned, everything possible done before the car arrives. Southwest Airlines built a whole low-cost empire on the same move: the famous "ten-minute turn," getting a plane back in the air while rivals took an hour, by prepping everything before the aircraft reached the gate. Same tool, different track: a plane earns nothing on the ground, and a machine earns nothing mid-changeover.
A crew switching between jobs is a changeover; so is an agency switching between clients. The fix is identical to the racetrack: stage everything the night before — materials picked, site notes read, briefs and assets loaded — so the paid clock starts on real work. External setup is free money in any business that bills by the day. — A.P.
Convert internal setup to external. The cheapest downtime is the setup you finished before the machine ever stopped.
- The WIP limit is the point of Kanban, not the board. Cap work-in-progress and hidden problems surface at once.
- Starting less finishes more. Juggling twelve completes none; finishing three beats it every time.
- Changeover cost silently sets your strategy. Slow switching forces big batches and kills flexibility.
- Prep before the clock starts. Converting internal setup to external recovers more time than raw speed.
KPIs — the Vital Few
A KPI translates an objective into a number that can be tracked over time, giving a team a shared, checkable definition of whether it is winning. Well-chosen indicators align effort: when everyone can see the same measure moving, coordination improves without constant intervention. The discipline of KPIs is the discipline of choosing the small number of measures that actually matter and resisting the temptation to track everything.
KPIs are commonly distinguished as lagging or leading. A lagging indicator measures an outcome that has already happened — last quarter's revenue, this month's defect count. A leading indicator measures something that predicts a future outcome and can still be influenced. Lagging indicators tell you whether you succeeded; leading indicators tell you whether you are about to. A dashboard built only of lagging indicators is a rear-view mirror.
The most useful KPI has three properties: it is tied directly to an objective that matters, it can be influenced by the people held accountable for it, and it is paired with a counter-measure that prevents it being satisfied at the expense of something unmeasured. An indicator missing any of these three either misleads, demoralises, or gets gamed.
Netflix famously obsessed over a single leading indicator — retention — over vanity metrics like sign-ups, reasoning that a customer who stays is the only one who was truly served. That's a well-chosen KPI. The opposite lesson comes from Wells Fargo, where a KPI on new accounts opened, pushed hard with bonuses and no counter-metric, drove staff to open millions of fraudulent accounts. Same tool, opposite outcomes: a KPI tied to genuine value guides a company; a KPI tied to a number with no guardrail detonates one.
My vital few in a service business: quote-to-close rate, on-time completion, and rework rate — the third existing purely to keep the first two honest. Everything else is context. The discipline isn't picking the three; it's saying no to the fourth, fifth, and fortieth. — A.P.
Measure everything that matters. Choose the vital few that predict success, pair each with a counter-metric, and leave the rest off.
OEE — the Honest Composite
OEE condenses three distinct questions into one number. Availability asks: of the time the process was scheduled to run, how much did it actually run, rather than sitting stopped? Performance asks: while running, did it run at its intended speed? Quality asks: of what it produced, how much was good on the first pass? Multiplying the three yields a single percentage that captures the true effectiveness of the asset.
The value of the composite is that it prevents the local optimisation Chapter I warned against. A process can be made to look busy by running fast (high performance) while producing scrap (low quality), or by running constantly (high availability) at reduced speed. Because OEE multiplies the three, no single factor can be inflated to disguise a failure in another. It is a metric structurally resistant to the gaming that afflicts single measures.
OEE is therefore best used as a diagnostic on the constraining step of a process, where every lost minute is a minute lost to the whole system. Applied there, it directs improvement precisely at the three ways an asset fails to deliver. Applied indiscriminately across every step, it can reward exactly the overproduction the rest of the toolkit exists to prevent.
Eliyahu Goldratt's business novel The Goal — required reading in operations for forty years — makes exactly this point through a plant manager saving his factory. His breakthrough: an hour lost at the bottleneck is an hour lost for the whole plant, while an hour saved at a non-bottleneck is a mirage. Teams there were producing beautiful efficiency numbers on machines that didn't matter, piling up inventory in front of the one that did. The warning in one line: a high score on the wrong machine isn't productivity — it's expensive, well-measured waste.
OEE translates cleanly out of the factory if you keep the multiplication: crew utilisation × schedule adherence × first-pass quality. Three pretty-good numbers still multiply into a mediocre one — and the same trap applies: chase it only on your constraint crew or team. A busy non-bottleneck is just well-documented waiting. — A.P.
Chase OEE on the bottleneck, and only there. Everywhere else, a high score may just be efficient waste.
- A KPI is a metric people will game. Everything in Chapter II applies here with a bonus cheque attached.
- If everything is key, nothing is. Pick the vital few leading indicators; leave the rest off the board.
- OEE multiplies, so it can't flatter. Three 90%s make 73% — the honesty is in the arithmetic.
- Chase OEE only on the bottleneck. Elsewhere a high score is overproduction with excellent paperwork.