From a formula in (x, t) to a moving point
You climbed this far by solving the heat equation — separation of variables, the heat kernel, Fourier modes, Green's functions. This new rung asks you to step back and look at the equation from a height. The single move that organizes all of parabolic theory is this: stop reading u(x, t) as a surface over the (x, t)-plane, and start reading it as a moving profile. At each frozen instant t you hold a whole spatial shape, the function u(., t) of x; as t advances, that shape drifts. The solution is no longer a surface — it is a path, a single point gliding through a space whose every element is an entire spatial profile.
Why bother re-reading something you can already solve? Because the explicit tools you mastered only work in lucky geometries — separation of variables needs a separable domain, the heat kernel needs the whole line or a box. The moving-point view costs you the explicit formula but buys you something far larger: a theory that handles any reasonable domain, variable coefficients, and even nonlinear terms, all at once. It turns a parabolic equation into a problem you already understand from a much earlier rung — an ordinary differential equation — only now living in an infinite-dimensional space.
Splitting time off from space
Look closely at what makes the heat equation u_t = k u_xx parabolic rather than elliptic or hyperbolic. It carries exactly one time derivative, to first order, while the spatial part is a second-order elliptic operator — a diffusion-like law that pulls every profile toward its local average. That lopsidedness, one privileged time direction plus a smoothing spatial operator, is the whole personality of the equation, and it is precisely what lets you peel time apart from space. The classification you learned earlier (the discriminant B^2 - A C) was the algebraic face of this; the moving-profile picture is its dynamical face.
Write the general parabolic equation as u_t + L u = f, where L is the elliptic spatial part. Now bundle everything spatial — the derivatives in x, and the boundary conditions — into one object A, an operator that acts on spatial profiles. For the plain heat equation with zero boundary data, A is the negative Laplacian, A = -Laplacian, but L could be a general elliptic operator with variable coefficients. The moment you do this, the PDE reads du/dt = -A u: the rate of change of the moving profile equals minus A applied to wherever the profile currently sits. This is the abstract Cauchy problem, the central object of the whole rung.
PDE in (x, t) abstract ODE in a function space ------------------ -------------------------------- u_t = k u_xx du/dt = -A u u = 0 on the boundary ...folded into the domain D(A) u(x, 0) = u_0(x) u(0) = u_0 u(., t) = a single moving point in the space X = L^2(Omega) A = -k * (d^2/dx^2), the spatial law of motion
Where the boundary conditions go: the domain of A
Here is the subtle part, and the place beginners stumble. When you fold space into A, the boundary conditions do not disappear — they migrate into the domain of A, the set of profiles A is even allowed to act on. The operator A = -Laplacian is not fully specified until you say which functions it eats. For the heat problem with zero data on the wall, A is the negative Laplacian restricted to profiles that are smooth enough (a Sobolev condition) and vanish on the boundary. That domain, written D(A), is not bookkeeping you can skip; it carries the boundary condition, and a different boundary condition is literally a different operator A.
And A is unbounded. That word means it cannot act on every profile in the space — feed it a jagged, non-differentiable shape and -Laplacian of it is meaningless. So D(A) is only a dense slice of the full space X: the smooth-enough profiles, dense but far from everything. This is exactly why du/dt = -A u is not the toy linear ODE it pretends to be. You cannot just declare the solution to be e^(-tA) u_0 as a power series and walk away, the way you would for a matrix A — an unbounded operator has no convergent exponential series. Making honest sense of e^(-tA) for such an A is the entire content of the next guides; the whole reason semigroup theory exists is to give that symbol a rigorous meaning.
The natural home of the path: Bochner spaces
If the solution is a curve t goes to u(t) through a function space, you need a place to measure such curves — a notion of "how big is this path, over the whole time interval?" That place is a Bochner space. It is the most natural generalization of the L^p spaces you already know: instead of integrating the size of a number, you integrate the spatial norm of a profile. The Bochner space L^2(0, T; X) collects maps u from the time interval (0, T) into the function space X for which the integral from 0 to T of (norm of u(t) in X)^2 dt is finite. It is the ordinary L^2 definition, with the absolute value swapped for the X-norm.
For a parabolic problem the right home is sharper than just X = L^2. The standard solution class asks the profile to live, most of the time, in the Sobolev space H^1_0 of finite-energy shapes that vanish on the wall — written u in L^2(0, T; H^1_0) — and its time-derivative u' to lie in the weaker dual space, u' in L^2(0, T; H^(-1)). The first clause says the profile keeps finite gradient energy as it moves; the second says its velocity is square-integrable in time, measured loosely. Together they are exactly strong enough to make u a continuous curve in time with values in L^2 — the path has no jumps. This is the precise function class in which guide 2 will construct a weak solution.
Why this rephrasing is honest — and why direction matters
Be honest about what the recasting does and does not buy you. It does not make the equation easy by sleight of hand — all the genuine difficulty has simply been relocated into a single object, the unbounded operator A and its domain. Every hard question about the PDE becomes a question about A: does the curve exist for every starting profile (does A generate a flow)? Is it unique? Does it stay bounded, or decay? The payoff is that these are now questions you can attack with one unified machinery — the spectrum and resolvent of A — instead of a different ad-hoc trick for every geometry. That trade, difficulty-concentrated-into-A, is the real content of the abstract viewpoint.
And one feature of the abstract ODE is no mere technicality — it is the deepest fact about parabolic evolution. The flow runs forward only. You may freely solve du/dt = -A u for t > 0, but you may not run it backward: the backward heat equation is ill-posed, in the strong sense that arbitrarily tiny changes in the data can produce arbitrarily huge changes in the answer. Diffusion smooths and so destroys fine information going forward; with the fine structure gone, nothing pins down where it came from. This is why the solution operator e^(-tA) will form only a semigroup — a C_0 semigroup, not a group — there is no e^(+tA) to undo it. The arrow of time is built into the equation, not added by hand.
So where does this leave you, standing at the foot of the rung? You now see the heat equation, and every parabolic equation, as one moving point obeying du/dt = -A u in a Bochner-measured path space, with all of space and its boundary conditions packed into one unbounded operator A. The remaining four guides simply interrogate A in turn: how to prove the curve exists (energy estimates and the maximum principle, guide 2), how its smoothing erases roughness at once (guide 3), how to build the flow e^(-tA) and its generator (guide 4), and how to add forcing and read off the long-time fate (guide 5). One change of viewpoint, four payoffs.