Meeting the delta on its own terms
The last guide introduced the Dirac delta in one line — < delta, phi > = phi(0) — and then walked on. Let us stop and really inhabit that rule, because almost everything in this rung orbits it. The delta is not a tall thin spike, no matter how often it is drawn that way. It is a measuring instrument: hand it any smooth probe phi and it reports back exactly one number, the height of that probe at the origin. It ignores phi everywhere else; it ignores phi's slope, its width, its area. All it ever says is 'your probe is this tall at 0'.
Why insist there is no function behind it? Suppose some honest function g had < g, phi > = integral of g(x) phi(x) dx equal to phi(0) for every probe. Then take probes that are zero at the origin but bulge up somewhere nearby; g would have to ignore all that bulk and answer 0, which forces g to be zero away from the origin. But a function that is zero everywhere except a single point integrates to nothing against everything. There is no escape: the delta is genuinely a distribution, a creature of the larger world the previous guide built, and not a disguised function. The picture of a spike is a useful crutch, but the rule is the truth.
The one legal move: throw the derivative onto the probe
Now the centrepiece. We want to differentiate objects that have no honest derivative — a corner, a jump, the delta itself. There is no 'inside' of a distribution to reach into, so we are stuck with the reflex the previous guide drilled: whatever we want to do to T, we instead do to the probe phi and let the definition carry it across. For the derivative, the bridge is integration by parts. For an honest differentiable function f, integrate f'(x) phi(x) and move the prime onto phi: the compact support of phi kills the boundary terms, leaving < f', phi > = - < f, phi' >. Both sides are now just rules acting on probes.
That identity, true for any smooth f, is what we promote into a definition. The distributional derivative T' of any distribution T is the rule defined by < T', phi > = - < T, phi' >. Read it slowly: to differentiate T, you do not touch T at all — you differentiate the probe, evaluate the old rule, and flip the sign. Because phi is a test function it is infinitely smooth, so phi' is always available; the minus sign is the only fee. This single formula gives every distribution a first derivative, hence a second, hence derivatives of every order, forever, with no corner or jump able to stop it. The bottleneck that opened this whole rung is gone.
The example everyone must see: a step differentiates into a delta
Let us run the machine on the cleanest possible jump. The Heaviside step H(x) is 0 for x < 0 and 1 for x > 0 — a switch that flips on at the origin. Classically it has no derivative at 0; the slope there is undefined, the very situation that broke us earlier. As a distribution, though, H is perfectly fine (it is locally integrable), so it has a distributional derivative, and the definition tells us exactly what it is. We just compute < H', phi > = - < H, phi' > and see what rule comes out.
Heaviside H(x) = 0 for x<0, 1 for x>0
< H', phi > = - < H, phi' > (definition)
= - integral over all x of H(x) phi'(x) dx
= - integral from 0 to infinity of phi'(x) dx (H = 1 there, 0 elsewhere)
= - [ phi(infinity) - phi(0) ]
= - [ 0 - phi(0) ] (phi has compact support)
= phi(0)
= < delta, phi >
Conclusion: H' = deltaStare at the result: H' = delta. The derivative of a switch flipping from 0 to 1 is a delta sitting precisely where the jump is, and its 'weight' is exactly the size of the jump. This is the rule of thumb to carry everywhere — differentiating a jump of height a produces a times a delta at the jump. It is no longer a heuristic or a physicist's shorthand; we proved it from the definition with nothing but integration by parts and compact support. The corner-and-jump catastrophe that motivated guide 1 has been turned into a clean, computable fact.
Two follow-ups deepen the picture. First, differentiate again: delta' is a legitimate distribution too, with < delta', phi > = - phi'(0) — it reads off the slope of the probe at the origin, not its height. So even the delta has a whole tower of derivatives. Second, a warning that saves real grief: if a function has a corner (a kink, like |x|, where the slope jumps but the value does not), its first distributional derivative is an ordinary bounded function (a step), and only the second derivative produces a delta. Match the order of the derivative to the order of the singularity — a jump in the value shows up one derivative sooner than a jump in the slope.
How to differentiate a distribution, step by step
The Heaviside calculation is a template, and almost every distributional differentiation you will ever do follows the same five beats. The skill is not cleverness; it is discipline — never reach inside T, always act on phi, and account honestly for the boundary terms (which vanish precisely because phi is compactly supported).
- Write the target as a rule on probes: pin down < T, phi > as a concrete integral (or evaluation, if T is already a delta-type object).
- Apply the definition: replace < T', phi > by - < T, phi' >, moving the derivative off T and onto the smooth probe.
- Integrate by parts where T is a function, splitting the integral at any jump or corner so each piece is smooth.
- Discard the boundary terms at infinity — phi has compact support, so phi and all its derivatives are zero out there — and collect the leftover boundary terms at the jumps.
- Read off the answer as a new rule: identify the resulting functional with a known distribution (an ordinary function plus, perhaps, deltas weighted by the jumps).
Run this on a sawtooth, a triangle wave, or any piecewise-smooth signal and the output is always the same shape: the classical derivative wherever the function is smooth, plus a delta at every jump scaled by that jump's height. That decomposition — a tame function part plus a sum of spikes pinned to the discontinuities — is the everyday face of the distributional derivative, and it is exactly what lets a distributional solution of a PDE carry a shock or a sharp front without anyone needing to differentiate the impossible.
Where this lands: kernels, sources, and honest small print
The reason this rung devotes so much care to the delta is that it is the natural language of a point source. When we model a sharp blow at one spot and one instant — a hammer tap on a string, a unit of heat injected at a single point — the source term is a delta, and the solution that answers it is a fundamental solution. The heat kernel, for instance, is precisely the temperature that develops when you start from a delta of heat at the origin: it is the response of the heat equation to a point source, a distribution in, a smooth function out. That 'response to a delta' framing is the whole subject of guide 5 in this rung, made rigorous.
Notice the beautiful asymmetry hiding here, because it is honest physics. Feed the heat equation a delta and an instant later you get a perfectly smooth bell curve: diffusion smooths instantly and propagates at infinite speed, so the worst possible initial data becomes infinitely differentiable at once. The wave equation does the opposite — it does not smooth, and it propagates at finite speed, so a delta launched into it stays sharp and rides outward on a front. Same delta source, utterly different responses, dictated by the type of the equation. The distributional framework handles both without flinching; it does not pretend the physics is uniform.
And one last piece of honest small print, repeated because it matters most exactly here. The distributional derivative is a triumph, but it does not make everything legal. You still cannot multiply two distributions, so you cannot write 'delta times delta' or differentiate H and then square it inside a nonlinear term and expect meaning — this is the genuine wall that makes nonlinear PDEs, shocks, and the well-posedness questions on this ladder so much harder than linear ones. Differentiation became free; multiplication did not. Keeping those two facts side by side is what separates a fluent user of distributions from someone who has merely learned the delta's party trick.