Why the old probes block the Fourier transform
By now distributions feel solid: a rough object is whatever it does when averaged against smooth probes, and you can differentiate it as often as you like by handing the derivative over to the probe. We want one more power — the Fourier transform, the tool that turns u_xx into a clean multiplication and so unlocks the heat, wave, and Laplace equations in one stroke. The natural plan is to repeat the rung's whole reflex: define the Fourier transform of a distribution by throwing the transform onto the test function. The plan is right. But the probes we built in guide 2 quietly sabotage it.
Here is the snag, and it is worth feeling precisely. A test function in guide 2 was a bump: smooth, and exactly zero outside a finite window. That compact support was a gift for derivatives, because the boundary terms in integration by parts sat where the probe was flatly zero. But the Fourier transform of a compactly supported bump is never again compactly supported — squeeze a function into a tiny window and its transform smears out across all of frequency space and never returns to zero. So if we throw the transform onto a bump, the result leaves the space of bumps. The transform won't stay home, and a definition that escapes its own room is no definition at all.
A roomier probe: the Schwartz space
The replacement probe is a Schwartz function, and the space of them is the Schwartz space, written S. A Schwartz function phi is still infinitely smooth, but instead of switching off completely outside a window it is allowed to spread everywhere — provided it dies away faster than any power of x as you go out to infinity, and so does every one of its derivatives. The slogan is 'rapidly decreasing': not just phi(x) but x^2 phi(x), x^100 phi'(x), any power times any derivative, all still go to zero at infinity. The model citizen is the Gaussian phi(x) = e^(-x^2), smooth everywhere and crushed to nothing exponentially fast on both sides.
Why is this the right room? Because the Schwartz space is closed under the Fourier transform: feed it a Schwartz function and you get another Schwartz function back, with no leakage at all. The reason is a beautiful symmetry the transform has — smoothness on one side becomes fast decay on the other, and fast decay becomes smoothness — so a function that is both smooth and fast-decaying is mapped to one that is again both. The Gaussian shows it off perfectly: the Fourier transform of e^(-x^2) is, up to constants, another Gaussian. The probe lands exactly where it started. Now the rung's reflex can work, because throwing the transform onto a Schwartz probe leaves us still inside S.
A distribution that is allowed to act on these roomier probes is called a tempered distribution. The word 'tempered' means 'tamed at infinity': because a Schwartz probe only decays fast rather than vanishing outright, the object it tests must not grow too violently far away, or the pairing integral would blow up. So tempered distributions are a slightly smaller club than the wild distributions of guide 2 — every tempered distribution is a distribution, but not every distribution is tempered. The price of fast-decaying probes is a mild growth bound on what they can probe; the reward, as we are about to see, is the Fourier transform.
Transferring the transform onto the probe
Now the definition, and it follows the exact pattern you used for the derivative in guide 3. For an ordinary, nicely decaying function f there is an old identity — a consequence of swapping the order of two integrals — that the integral of (Fourier transform of f) times phi equals the integral of f times (Fourier transform of phi). In probe-pairing notation this reads < F f, phi > = < f, F phi >. The left side asks for the transform of f directly; the right side never touches f's transform at all — it pushes the whole transform onto the probe phi, which, being Schwartz, can take it and stay Schwartz.
So we promote that identity from a theorem about functions into a definition for distributions. The Fourier transform of a tempered distribution T is the new tempered distribution F T whose action on any Schwartz probe is decreed to be < F T, phi > = < T, F phi >. Read it slowly: to find out what the transform of T does to a probe, you do not transform T — you transform the probe and then pair the original T against it. The whole operation is carried out in a place where it is perfectly safe, on the smooth fast-decaying phi, and T itself is never asked to do anything it cannot. This is the same self-denial as always: you cannot reach inside T, so you act on the probe and let the identity carry the operation across.
What the transform now bites on
Watch the definition pay off on the object that started this whole rung. What is the Fourier transform of the Dirac delta? By the rule, < F delta, phi > = < delta, F phi >, and the delta just reads off the value of its argument at the origin, so this equals (F phi)(0). But the Fourier transform of phi evaluated at frequency zero is simply the integral of phi over all of x — that is, the pairing of the constant function 1 against phi. So < F delta, phi > = < 1, phi > for every probe, which means F delta = 1. The transform of a single infinite spike at the origin is the flat constant 1: a perfectly localized thing in position becomes perfectly spread out in frequency, exactly as the uncertainty slogan predicted.
Run it the other way and it is just as telling. The transform of the constant function 1 is the delta (up to a 2-pi convention) — a flat, never-decaying constant, which the classical transform could not touch at all because its integral diverges, becomes a clean spike at frequency zero. And the Heaviside step H, the jump from 0 to 1, has a tempered transform too, built from a delta at zero plus a principal value of 1/(i times frequency) — the careful 'symmetric limit' reading of the singular 1/x that distributions make rigorous. None of these — a constant, a step — has a classical Fourier transform, yet all three now sit comfortably inside the theory.
definition : < F T, phi > = < T, F phi > (move F onto the Schwartz probe) F[ delta ] = 1 (spike at 0 -> flat constant) F[ 1 ] = 2*pi * delta (flat constant -> spike at 0) F[ H(x) ] = pi*delta(k) + p.v. 1/(i k) (step -> delta + principal value) derivative rule : F[ u' ] = (i k) * F[ u ] (d/dx becomes multiply by i k) convolution rule : F[ f * g ] = F[ f ] * F[ g ] (convolution becomes product)
The two rules on the last lines are why we bothered with all this. The differentiation rule says the transform turns d/dx into multiplication by i times the frequency, and it now holds for any tempered distribution, no smoothness required. That is the magic that flattens a PDE: u_t = k u_xx becomes, after transforming in x, an ordinary differential equation in t for each frequency, because u_xx turns into minus-frequency-squared times the transform. And the convolution rule turns the messy integral of a convolution into a plain product of transforms — which is exactly how, in guide 5, applying an operator to a fundamental solution will become multiplying by its symbol and then dividing back.
Honest limits, and the bridge to fundamental solutions
Keep the caveats from guide 2 firmly in view, because moving to the Fourier side does not lift them. You still cannot freely multiply two distributions, and the transform makes this vivid: by the convolution rule, multiplying on one side is convolution on the other, so an illegal product of two transforms would correspond to a convolution that may not converge — the product problem reappears in a new costume. And tempered-ness is a real restriction: a function that grows like e^(x^2) is a perfectly good distribution but is not tempered, because it overwhelms even a Schwartz probe's fast decay, so it has no tempered Fourier transform. The Fourier world is powerful but it is not the whole world; you pay for the transform with a growth bound.
Step back and see how far one reflex has carried us. We swapped compactly supported bumps for the fast-decaying probes of the Schwartz space, named the objects that ride on them the tempered distributions, and then defined the Fourier transform of any one of them by the single line < F T, phi > = < T, F phi >. In return the transform finally bites on the delta, the constant, and the step, and it carries the two rules — derivative-to-multiplication and convolution-to-product — that flatten a constant-coefficient PDE into algebra. Every piece needed for the rung's payoff is now on the table.
The last guide of this rung cashes it all in. To build a fundamental solution — the response E of an operator L to a unit point source, the precise meaning of L E = delta — you now transform both sides: F delta is the constant 1, the operator L becomes multiplication by its symbol, and solving collapses to dividing 1 by that symbol and transforming back. The infinite spike that ordinary functions could not hold, the response that classical calculus could not differentiate into being, the transform that classically could not even look at it — all three repairs of this rung meet in that one calculation. Guide 5 makes it rigorous.