1 Relativistic Quantum Mechanics

In order to describe the dynamics of particles involved in high-energy collisions we must be able to combine the theory of phenomena occurring at the smallest scales, i.e.  quantum mechanics (QM), with the description of particles moving close to the speed of light, i.e. special relativity. To do this we must develop wave equations which are relativistically invariant (i.e. invariant under Lorentz transformations). In this section we will derive relativistic equations of motion for scalar particles (spin-0) and particles with spin-12\tfrac{1}{2}.

The Standard Model Lagrangian in popular form reads

ℒ=−14⁢Fμ⁢ν⁢Fμ⁢ν+i⁢ψ¯⁢γ⋅D⁢ψ+ψ¯i⁢yi⁢j⁢ψj⁢ϕ+|Dμ⁢ϕ|2−V⁢(ϕ)\displaystyle\begin{split}\mathcal{L}=&-\frac{1}{4}F_{\mu\nu}F^{\mu\nu}\\ &+i\bar{\psi}\gamma\cdot D\psi\\ &+\bar{\psi}_{i}y_{ij}\psi_{j}\phi\\ &+|D_{\mu}\phi|^{2}-V(\phi)\end{split} (1)

where the first line accounts for the self-interactions of the pure gauge fields, the second the free Dirac fields and their interactions with the gauge fields, the third line the interaction between the fermions and the Higgs field (i.e. the generation of masses to the fermions) and finally the fourth line the Higgs boson field and its interaction with the gauge fields. The focus of this brief course is to explain the notation contained in the two middle lines, and explore some of the complexity of the fields ψ\psi and the operator DD.

1.1 The Klein-Gordon Equation

Consider first the Hamiltonian for a particle in classical (non-relativistic) mechanics

E=p→22⁢m+V⁢(x→).\displaystyle E=\frac{\vec{p}^{2}}{2m}+V(\vec{x})\,. (2)

To convert this into a wave equation for QM, we use the identifications E→i⁢∂tE\to i\partial_{t} and p→→−i⁢∇→\vec{p}\to-i\vec{\nabla}, so that a plane-wave solution

ϕ⁢(t,x→)=A⁢e−i⁢(E⁢t−p→⋅x→)=A⁢e−i⁢p⋅x\displaystyle\phi(t,\vec{x})=A{\rm e}^{-i(Et-\vec{p}\cdot\vec{x})}=A{\rm e}^{-% ip\cdot x} (3)

has the energy-momentum relation given in (2). Applied to a general wavefunction ϕ\phi, a linear superposition of plane waves, this gives

i⁢∂tϕ⁢(t,x→)=(−12⁢m⁢∇→2+V⁢(x→))⁢ϕ⁢(t,x→)=H^⁢ϕ⁢(t,x→),\displaystyle i\partial_{t}\phi(t,\vec{x})=\bigg{(}-\frac{1}{2m}\vec{\nabla}^{% 2}+V(\vec{x})\bigg{)}\phi(t,\vec{x})=\hat{H}\phi(t,\vec{x})\,, (4)

where H^\hat{H} is the so-called Hamiltonian operator. We recognise this as the Schrödinger Equation, the cornerstone of QM. (4) cannot be relativistically invariant because time appears only through a first-order derivative on the left-hand side while space appears as a second-order derivative on the right-hand side. Yet we know that if we make a Lorentz transformation, this would mix the x→\vec{x} and tt components and therefore their derivatives will have to arise with the same orders.

The problem with the Schrödinger Equation arose because we started from a non-relativistic energy-momentum relation. Let us then start from the relativistic equation for energy. For a particle with 4-momentum pμ=(E,p→)p^{\mu}=(E,\vec{p}) and mass mm,

E2=m2+p→2.\displaystyle E^{2}=m^{2}+\vec{p}^{2}\,. (5)

Again we convert this to an operator equation by setting pμ=i⁢∂μp_{\mu}=i\partial_{\mu} so that the corresponding wave equation for an arbitrary scalar wavefunction ϕ⁢(x→,t)\phi(\vec{x},t) gives

(∂t2−∇2+m2)⁢ϕ⁢(t,x→)=(∂μ∂μ+m2)⁢ϕ⁢(x)=0,\displaystyle(\partial_{t}^{2}-\nabla^{2}+m^{2})\phi(t,\vec{x})=(\partial_{\mu% }\partial^{\mu}+m^{2})\phi(x)=0\,, (6)

where we have introduced the four-vector xμ=(t,x→)x^{\mu}=(t,\vec{x}). This is the Klein-Gordon (KG) equation which is the equation of motion for a free scalar field. We can explicitly check that this is indeed Lorentz invariant.

Under a Lorentz transformation the contravariant four-vector xμx^{\mu} and covariant ∂μ\partial_{\mu} transform as

xμ→(x′)μ=Λνμ⁢xν,∂μ→(∂′)μ=Λμρ⁢∂ρ.\displaystyle x^{\mu}\to(x^{\prime})^{\mu}=\Lambda^{\mu}_{\phantom{\mu}\nu}x^{% \nu}\,,\qquad\partial_{\mu}\to(\partial^{\prime})_{\mu}=\Lambda_{\mu}^{% \phantom{\mu}\rho}\partial_{\rho}\,. (7)

The Lorentz transformations have the special property that

Λνμ⁢Λμρ=δνρ=δνρ.\displaystyle\Lambda^{\mu}_{\phantom{\mu}\nu}\Lambda_{\mu}^{\phantom{\mu}\rho}% =\delta_{\nu}^{\phantom{\nu}\rho}=\delta^{\rho}_{\phantom{\rho}\nu}\,. (8)

This equation says that the inverse of transformation Λνμ\Lambda^{\mu}_{\phantom{\mu}\nu} is given by the transformation Λμρ\Lambda_{\mu}^{\phantom{\mu}\rho}.

The field ϕ\phi is a scalar, i.e. it has the transformation property

ϕ⁢(x)→ϕ′⁢(x′)=ϕ′⁢(Λ⁢x)=ϕ⁢(x)\displaystyle\phi(x)\to\phi^{\prime}(x^{\prime})=\phi^{\prime}(\Lambda x)=\phi% (x) (9)

Therefore, in the primed system one obtains, using the aforementioned properties of the Lorentz transformations

[(∂′)μ⁡(∂′)μ+m2]⁢ϕ′⁢(x′)=[Λμρ⁢∂ρΛσμ⁢∂σ+m2]⁢ϕ′⁢(Λ⁢x)=[∂ρ∂ρ+m2]⁢ϕ⁢(x)=0\displaystyle\Big{[}(\partial^{\prime})_{\mu}(\partial^{\prime})^{\mu}+m^{2}% \Big{]}\phi^{\prime}(x^{\prime})=\Big{[}\Lambda_{\mu}^{\phantom{\mu}\rho}% \partial_{\rho}\Lambda^{\mu}_{\phantom{\mu}\sigma}\partial^{\sigma}+m^{2}\Big{% ]}\phi^{\prime}(\Lambda x)=\Big{[}\partial_{\rho}\partial^{\rho}+m^{2}\Big{]}% \phi(x)=0 (10)

and the equation still holds.

1.2 The Dirac Equation

The KG equation admits negative-energy solutions, because the energy EE appearing in the plane-wave in (3) can have the two values ±p→2+m2\pm\sqrt{\vec{p}^{2}+m^{2}}. This originates from the energy-momentum relation E2=p→2+m2E^{2}=\vec{p}^{2}+m^{2}, (5). Dirac sought to find an alternative relativistic equation which was linear in E∼∂tE\sim\partial_{t} like the Schrödinger equation since looks like it would remove the negative energy solution. If the equation is linear in ∂t\partial_{t}, it must also be linear in ∇\nabla if it is to be invariant under Lorentz transformations.

1.2.1 An alternative wave equation

We therefore start with the general form

i⁢∂tψ⁢(t,x→)=(−i⁢α→⋅∇→+β⁢m)⁢ψ⁢(t,x→).\displaystyle i\partial_{t}\psi(t,\vec{x})=(-i\vec{\alpha}\cdot\vec{\nabla}+% \beta m)\psi(t,\vec{x})\,. (11)

We also want the solution to this equation to follow the energy relation (5). To implement this constraint, we square (11)

−∂t2ψ⁢(t,x→)\displaystyle-\partial_{t}^{2}\psi(t,\vec{x}) =i⁢∂t(−i⁢α→⋅∇→+β⁢m)⁢ψ⁢(t,x→)\displaystyle=i\partial_{t}\Big{(}-i\vec{\alpha}\cdot\vec{\nabla}+\beta m\Big{% )}\psi(t,\vec{x}) (12a)
=(−i⁢α→⋅∇→+β⁢m)2⁢ψ⁢(t,x→)\displaystyle=\Big{(}-i\vec{\alpha}\cdot\vec{\nabla}+\beta m\Big{)}^{2}\psi(t,% \vec{x}) (12b)
=(−αi⁢αj⁢∂i∂j−i⁢(β⁢αi+αi⁢β)⁢m⁢∂i+β2⁢m2)⁢ψ⁢(t,x→)\displaystyle=\Big{(}-\alpha^{i}\alpha^{j}\partial_{i}\partial_{j}-i(\beta% \alpha^{i}+\alpha^{i}\beta)m\partial_{i}+\beta^{2}m^{2}\Big{)}\psi(t,\vec{x}) =!(−∇2+m2)⁢ψ⁢(t,x→).\displaystyle\stackrel{{\scriptstyle!}}{{=}}\Big{(}-\nabla^{2}+m^{2}\Big{)}% \psi(t,\vec{x})\,. (12c)

The KG equations, which is just the energy-momentum relation, requires the equality marked with the exclamation mark. Therefore, for α→\vec{\alpha} and β\beta

αi⁢αj+αj⁢αi\displaystyle\alpha^{i}\alpha^{j}+\alpha^{j}\alpha^{i} ={αi,αj}=2⁢δi⁢j,\displaystyle=\{\alpha^{i},\alpha^{j}\}=2\delta^{ij}\,, (13a)
β⁢αi+αi⁢β\displaystyle\beta\alpha^{i}+\alpha^{i}\beta ={β,αi}=0,\displaystyle=\{\beta,\alpha^{i}\}=0\,, (13b)
β2\displaystyle\beta^{2} =1.\displaystyle=1\,. (13c)

Here we have defined the anti-commutator {A,B}=A⁢B+B⁢A\{A,B\}=AB+BA which is similar to the normal commutator [A,B]=A⁢B−B⁢A[A,B]=AB-BA. It is clearly not possible to satisfy these equations for numbers αi\alpha^{i} and β\beta. Instead, Dirac defined n×nn\times n for αi\alpha^{i} and β\beta and ψ\psi to be a column vector. One can show that (13) requires that

tr⁢αi=tr⁢β=0\displaystyle{\rm tr}\alpha^{i}={\rm tr}\beta=0 (14)

and that the eigenvalues of all matrices need to be ±1\pm 1. Therefore, nn needs to be even. The simplest case of n=2n=2 has only three independent matrices, the Pauli matrices

σx=(0110),σy=(0−ii0),σz=(100−1),\displaystyle\sigma_{x}=\begin{pmatrix}0&1\\ 1&0\end{pmatrix}\,,\qquad\sigma_{y}=\begin{pmatrix}0&-i\\ i&0\end{pmatrix}\,,\qquad\sigma_{z}=\begin{pmatrix}1&0\\ 0&-1\end{pmatrix}\,, (15)

which is not enough. Therefore, the simplest solution is n=4n=4, for example,

αi=(0σiσi0),β=(12×200−12×2).\displaystyle\alpha^{i}=\begin{pmatrix}0&\sigma^{i}\\ \sigma^{i}&0\end{pmatrix}\,,\qquad\beta=\begin{pmatrix}1_{2\times 2}&0\\ 0&-1_{2\times 2}\end{pmatrix}\,. (16)

Note that the σi\sigma^{i} and the 12×21_{2\times 2} here are both 2×22\times 2 matrices and that this block notation is just a shorthand notation. In the future, I will drop the 2×22\times 2 subscript. It is customary to collect these matrices as

γ0=β,γi=β⁢αi\displaystyle\gamma^{0}=\beta\,,\qquad\gamma^{i}=\beta\alpha^{i} (17)

and to define γμ=(γ0,γ→)\gamma^{\mu}=(\gamma^{0},\vec{\gamma}) as a new Lorentz vector where each component is a 4×44\times 4 matrix. The condition  (13) then becomes

{γμ,γν}=γμ⁢γν+γν⁢γμ=2⁢gμ⁢ν.\displaystyle\{\gamma^{\mu},\gamma^{\nu}\}=\gamma^{\mu}\gamma^{\nu}+\gamma^{% \nu}\gamma^{\mu}=2g^{\mu\nu}\,. (18)

With these matrices we can now write down the Dirac equation by multiplying (11) with γ0\gamma^{0}

(i⁢γμ⁢∂μ−m⁢ 14)⁢ψ⁢(t,x→)=(i⁢γ⋅∂−m)⁢ψ⁢(x)=0.\displaystyle(i\gamma^{\mu}\partial_{\mu}-m\ 1_{4})\psi(t,\vec{x})=(i\gamma% \cdot\partial-m)\psi(x)=0\,. (19)

In momentum space, i.e. setting ∂μ→−i⁢pμ\partial_{\mu}\to-ip_{\mu} via Fourier transformation, this becomes

(γ⋅p−m)⁢ψ~⁢(p)=0.\displaystyle(\gamma\cdot p-m)\tilde{\psi}(p)=0\,. (20)

These matrices are an example of a Clifford algebra. In fact, any set of matrices that fulfils (18) can be used to construct the Dirac equation. The representation in (16) is just an example, known as the Dirac representation. It is possible to find another representation by transforming

αi′=U⁢αi⁢U−1,andβ′=U⁢β⁢U−1,\displaystyle\alpha_{i}^{\prime}=U\alpha_{i}U^{-1}\,,\qquad\text{and}\qquad% \beta^{\prime}=U\beta U^{-1}\,, (21)

where UU is a unitary matrix.

We mentioned in passing that ψ⁢(x)\psi(x) is a column vector rather than a scalar. This means that it contains more than one degree of freedom. Dirac exploited this property to interpret his equation as the wave equation for spin-1/2 particles, fermions, which can be either spin-up or spin-down. The column vector ψ\psi is known as a Dirac spinor, and the matrices γμ\gamma^{\mu} operate on these Dirac spinors.

1.2.2 Negative energy solutions

Compare the Schrödinger equation (4) to the Dirac equation. This gives the Hamiltonian for a free fermion as

H=−i⁢α→⋅∇→+β⁢m.\displaystyle H=-i\vec{\alpha}\cdot\vec{\nabla}+\beta\ m\,. (22)

The trace of the Hamiltonian gives us the sum of the energy eigenvalues for all internal degrees of freedoms. Since α→\vec{\alpha} and β\beta assumed to be traceless, the eigenvalues of HH must sum to zero. Therefore, we still have negative energy solutions, just like the KG equation!

Dirac’s solution to this problem is called the Dirac sea, cf. Figure 1. The negative-energy states are real but the vacuum is defines as the having all of those states already filled. This solves the practical problems because observations rely on energy differences but leaves a vacuum with infinite negative charge and energy.

An energy level diagram that covers positive and negative energies. There is a gap between -m and +m with no levels. There are three scenarios depicted, from left to right: all negative energies are filled; all negative energies are filled and one positive; all negative energies except the highest are filled as well as one positive one
Figure 1: The energy levels in the Dirac sea picture. They must satisfy |E|>m|E|>m, but negative-energy states are allowed. The vacuum (left) is the state in which all negative-energy levels are filled.

Since it can be shown that fermions follow Pauli’s exclusion principle (spin-statistics theorem: particles that use anti-commutators rather than commutators have spin 12\tfrac{1}{2}), a positive-energy electron cannot fall into the negative-energy sea. It is however possible to excite on the negative-energy states into a positive-energy state, leaving behind a hole. This hole can be interpreted as a state with positive energy and positive charge, called a positron. Dirac predicted the existence of these particles in 1927 and they were experimentally confirmed in 1932.

A more consistent solution to this problem that also covers the KG equation was proposed later by Feynman and Stückelberg, based on considering the plane wave solution (3) which has a term E⋅tE\cdot t. Their idea is to consider a particle that has E>0E>0 to propagate forwards in time and a particle with E<0E<0 to propagate backwards in time.