Sunday, January 14, 2018
New Location
I have a new blog at github pages. All of the content of this blog has been migrated there. This blog is no longer being maintained.
Tuesday, September 15, 2015
Santalo Formula
Let $M$ be a simple Riemannian manifold with boundary $\partial M$. For $(x,v) \in SM$, let $\tau(x,v)$ denote the exit time of the geodesic starting at $x$ with tangent vector $v$, i.e. $\tau(x,v)$ is the (necessarily unique, and finite) time at which $\exp_x(tv) \in \partial M$.
We let $\partial_+(SM)$ denote the set
\[ \partial_+(SM) = \{ (x,v) \in SM \ | \ x \in \partial M, \langle v, \nu\rangle > 0 \} \]
where $\nu$ denotes the inward unit normal to $\partial M$ in $M$. The exponential map identifies $SM$ with the set
\[ \Omega = \{ (x,v,t) \in \partial_+(SM) \times \mathbf{R} \ | \ 0 \leq t \leq \tau(x,v) \}, \]
via $(x,v,t) \mapsto \exp_x(tv)$. Let $\Phi: \Omega \to M$ denote this diffeomorphism. Then we have, for all $f \in C^\infty(SM)$
\begin{align}
\int_{SM} f dvol(SM) &= \int_\Omega (\Phi^\ast f) (\Phi^\ast dvol(SM)) \\
&= \int_{\partial_+(SM)} \int_0^{\tau(x,v)} f(\phi_t(x,v)) \Phi^\ast dvol(SM).
\end{align}
Therefore, we can compute integrals of functions over $SM$ by integrating along geodesics, provided that we can cmopute $\Phi^\ast dvol(SM)$. This is the content of the Santalo formula.
Theorem (Santalo formula). For all $f \in C^\infty(SM)$, we have
\[ \int_{SM} f dvol(SM) = \int_{\partial_+(SM)} \int_0^{\tau(x,v)} f(\phi_t(x,v)) \langle v, \nu\rangle dt dvol(\partial(SM)) \]
Proof. Necessarily, we must have
\[ \Phi^\ast(dvol(SM)) = a(x,v) dt \wedge dvol(\partial(SM))), \]
for some function $a(x,v)$. The reason we can assume that $a$ is independent of $t$ is that $\Phi$ is defined via geodesic flow, and geodesic flow preserves the volume form on $SM$. To compute the factor $a(x,v)$, we just need to compute
\[ i_{\partial / \partial t} \Phi^\ast(dvol(SM)) = \Phi^\ast(i_{\Phi_\ast(\partial / \partial t)} dvol(SM)) \]
From the definition of $\Phi$, we have that $\Phi_\ast(\partial / \partial t)$ is the Reeb vector field on $SM$, i.e. the vector field generating geodesic flow. Therefore, $\Phi_\ast(\partial / \partial t)$ is equal, at a point $(x,v)$ to the horizontal lift of the vector $v$. Therefore, using the definition of the induced volume form on a hypersurface of a Riemannian manifold, we find
\[ i_{\partial / \partial t} \Phi^\ast(dvol(SM)) = \langle v, \nu \rangle dvol(\partial(SM)) \]
where $\nu$ is the inward pointing unit normal to $\partial(SM)$ in $SM$. This shows that $a(x,v) = \langle v, \nu \rangle$ and completes the proof.
We let $\partial_+(SM)$ denote the set
\[ \partial_+(SM) = \{ (x,v) \in SM \ | \ x \in \partial M, \langle v, \nu\rangle > 0 \} \]
where $\nu$ denotes the inward unit normal to $\partial M$ in $M$. The exponential map identifies $SM$ with the set
\[ \Omega = \{ (x,v,t) \in \partial_+(SM) \times \mathbf{R} \ | \ 0 \leq t \leq \tau(x,v) \}, \]
via $(x,v,t) \mapsto \exp_x(tv)$. Let $\Phi: \Omega \to M$ denote this diffeomorphism. Then we have, for all $f \in C^\infty(SM)$
\begin{align}
\int_{SM} f dvol(SM) &= \int_\Omega (\Phi^\ast f) (\Phi^\ast dvol(SM)) \\
&= \int_{\partial_+(SM)} \int_0^{\tau(x,v)} f(\phi_t(x,v)) \Phi^\ast dvol(SM).
\end{align}
Therefore, we can compute integrals of functions over $SM$ by integrating along geodesics, provided that we can cmopute $\Phi^\ast dvol(SM)$. This is the content of the Santalo formula.
Theorem (Santalo formula). For all $f \in C^\infty(SM)$, we have
\[ \int_{SM} f dvol(SM) = \int_{\partial_+(SM)} \int_0^{\tau(x,v)} f(\phi_t(x,v)) \langle v, \nu\rangle dt dvol(\partial(SM)) \]
Proof. Necessarily, we must have
\[ \Phi^\ast(dvol(SM)) = a(x,v) dt \wedge dvol(\partial(SM))), \]
for some function $a(x,v)$. The reason we can assume that $a$ is independent of $t$ is that $\Phi$ is defined via geodesic flow, and geodesic flow preserves the volume form on $SM$. To compute the factor $a(x,v)$, we just need to compute
\[ i_{\partial / \partial t} \Phi^\ast(dvol(SM)) = \Phi^\ast(i_{\Phi_\ast(\partial / \partial t)} dvol(SM)) \]
From the definition of $\Phi$, we have that $\Phi_\ast(\partial / \partial t)$ is the Reeb vector field on $SM$, i.e. the vector field generating geodesic flow. Therefore, $\Phi_\ast(\partial / \partial t)$ is equal, at a point $(x,v)$ to the horizontal lift of the vector $v$. Therefore, using the definition of the induced volume form on a hypersurface of a Riemannian manifold, we find
\[ i_{\partial / \partial t} \Phi^\ast(dvol(SM)) = \langle v, \nu \rangle dvol(\partial(SM)) \]
where $\nu$ is the inward pointing unit normal to $\partial(SM)$ in $SM$. This shows that $a(x,v) = \langle v, \nu \rangle$ and completes the proof.
Friday, September 11, 2015
The Index Form
Let $f: [0,T] \times (-\epsilon, \epsilon) \to M$ be a family of parametrized curves in a Riemannian manifold $(M, g)$. To simplify this calculation, we assume that $f(0,s) = p, f(T, s) = q$ for some $p,q \in M$ and all $s \in (-\epsilon, \epsilon)$. (This assumption is not necessary, but without it our variational formulae will have additional boundary terms.)
For convenience, set $\dot f = \partial f / \partial t$ and $f' = \partial f / \partial s$. For each $s \in (-\epsilon, \epsilon)$ we define the energy functional $E = E(s)$ to be
\[ E(s) = \frac{1}{2} \int_0^T |\dot f|^2 dt. \]
The first variation is
\begin{align}
\frac{dE}{ds} &= \int_0^T \langle \nabla_{f'} \dot f, \dot f \rangle dt \\\
&= \int_0^T \langle \nabla_{\dot f} f', \dot f \rangle dt \\\
&= -\int_0^T \langle f', \nabla_{\dot f}\dot f \rangle dt
\end{align}
Set $\gamma(t) := f(t,0)$ and $X(t) = f'(t)$ (thought of as a vector field supported on $\gamma$). Evaluating the above at $s=0$ we obtain
\[ \left.\frac{dE}{ds}\right|_{s=0} = -\int_0^T \langle X, \nabla_{\dot \gamma} \dot \gamma \rangle dt, \]
which shows immediately that
Theorem. $\gamma$ is a critical point of the energy functional if and only if $\nabla_{\dot \gamma} \dot \gamma = 0$.
The second variation is
\begin{align}
\frac{d^2 E}{ds^2}
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
+ \langle f', \nabla_{f'}\nabla_{\dot f}\dot f \rangle dt \\\
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
+ \langle f', \nabla_{\dot f}\nabla_{f'}\dot f \rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt \\\
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
- \langle \nabla_{\dot f}f', \nabla_{f'}\dot f \rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt \\\ &= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
- \langle \nabla_{\dot f} f', \nabla_{\dot f} f'\rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt
\end{align}
Assume now that $\gamma$ is a geodesic, i.e. $\nabla_{\dot \gamma} \dot \gamma = 0$. Then evaluating the above at $s=0$, we obtain
\[ \frac{d^2 E}{ds^2} = \int_0^T |\nabla_{\dot \gamma} X|^2 - \langle X, R(X, \dot \gamma) \dot \gamma \rangle dt. \]
Definition. Let $\gamma$ be a geodesic. The index form associated to variations $X,Y$ of $\gamma$ is
\begin{align} I(X,Y) &= \int_0^T \langle \nabla_{\dot \gamma} X, \nabla_{\dot \gamma} Y \rangle dt
- \langle Y, R(X, \dot \gamma) \dot \gamma \rangle \\\
&= -\int_0^T \langle Y, \nabla_{\dot \gamma}^2 X + R(X, \dot\gamma)\dot \gamma \rangle
\end{align}
It follows from symmetries of the Riemann tensor that $I(X,Y) = I(Y, X)$ and also $I(X,X) = E''$ as above.
Theorem. Suppose that $X$ is the infinitesimal variation of a family of affine geodesics about a fixed geodesic $\gamma$. Then
\[ \nabla_{\dot \gamma}^2 X + R(X, \dot\gamma)\dot\gamma = 0. \]
In particular, $I(X, -) = 0$.
Proof. Let $f(t,s)$ denote the family as above. By hypothesis, we have that $\nabla_{\dot f} \dot f = 0$ for all $s$, so that
\[ \nabla_{f'} \nabla_{\dot f} \dot f = 0. \]
Commuting the derivatives using the curvature tensor, we have
\[ 0 = \nabla_{\dot f} \nabla_{f'} \dot f + R(f', \dot f) \dot f. \]
Now use $\nabla_{\dot f} f' = \nabla_{f'} \dot f$ and evaluate at $s=0$ to obtain
\[ 0 = \nabla_{\dot \gamma}^2 X + R(X, \dot \gamma)\dot\gamma. \]
For convenience, set $\dot f = \partial f / \partial t$ and $f' = \partial f / \partial s$. For each $s \in (-\epsilon, \epsilon)$ we define the energy functional $E = E(s)$ to be
\[ E(s) = \frac{1}{2} \int_0^T |\dot f|^2 dt. \]
The first variation is
\begin{align}
\frac{dE}{ds} &= \int_0^T \langle \nabla_{f'} \dot f, \dot f \rangle dt \\\
&= \int_0^T \langle \nabla_{\dot f} f', \dot f \rangle dt \\\
&= -\int_0^T \langle f', \nabla_{\dot f}\dot f \rangle dt
\end{align}
Set $\gamma(t) := f(t,0)$ and $X(t) = f'(t)$ (thought of as a vector field supported on $\gamma$). Evaluating the above at $s=0$ we obtain
\[ \left.\frac{dE}{ds}\right|_{s=0} = -\int_0^T \langle X, \nabla_{\dot \gamma} \dot \gamma \rangle dt, \]
which shows immediately that
Theorem. $\gamma$ is a critical point of the energy functional if and only if $\nabla_{\dot \gamma} \dot \gamma = 0$.
The second variation is
\begin{align}
\frac{d^2 E}{ds^2}
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
+ \langle f', \nabla_{f'}\nabla_{\dot f}\dot f \rangle dt \\\
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
+ \langle f', \nabla_{\dot f}\nabla_{f'}\dot f \rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt \\\
&= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
- \langle \nabla_{\dot f}f', \nabla_{f'}\dot f \rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt \\\ &= -\int_0^T \langle \nabla_{f'}f', \nabla_{\dot f}\dot f \rangle
- \langle \nabla_{\dot f} f', \nabla_{\dot f} f'\rangle dt
+ \langle f', R(f', \dot f)\dot f \rangle dt
\end{align}
Assume now that $\gamma$ is a geodesic, i.e. $\nabla_{\dot \gamma} \dot \gamma = 0$. Then evaluating the above at $s=0$, we obtain
\[ \frac{d^2 E}{ds^2} = \int_0^T |\nabla_{\dot \gamma} X|^2 - \langle X, R(X, \dot \gamma) \dot \gamma \rangle dt. \]
Definition. Let $\gamma$ be a geodesic. The index form associated to variations $X,Y$ of $\gamma$ is
\begin{align} I(X,Y) &= \int_0^T \langle \nabla_{\dot \gamma} X, \nabla_{\dot \gamma} Y \rangle dt
- \langle Y, R(X, \dot \gamma) \dot \gamma \rangle \\\
&= -\int_0^T \langle Y, \nabla_{\dot \gamma}^2 X + R(X, \dot\gamma)\dot \gamma \rangle
\end{align}
It follows from symmetries of the Riemann tensor that $I(X,Y) = I(Y, X)$ and also $I(X,X) = E''$ as above.
Theorem. Suppose that $X$ is the infinitesimal variation of a family of affine geodesics about a fixed geodesic $\gamma$. Then
\[ \nabla_{\dot \gamma}^2 X + R(X, \dot\gamma)\dot\gamma = 0. \]
In particular, $I(X, -) = 0$.
Proof. Let $f(t,s)$ denote the family as above. By hypothesis, we have that $\nabla_{\dot f} \dot f = 0$ for all $s$, so that
\[ \nabla_{f'} \nabla_{\dot f} \dot f = 0. \]
Commuting the derivatives using the curvature tensor, we have
\[ 0 = \nabla_{\dot f} \nabla_{f'} \dot f + R(f', \dot f) \dot f. \]
Now use $\nabla_{\dot f} f' = \nabla_{f'} \dot f$ and evaluate at $s=0$ to obtain
\[ 0 = \nabla_{\dot \gamma}^2 X + R(X, \dot \gamma)\dot\gamma. \]
Thursday, September 3, 2015
Boundary Distance
Recently, I've been learning some topics related to machine learning, and especially manifold learning. These both fall under the general notion of inverse problems: given some mathematical object $X$ (it could be a function $f: A \to B$, or a Riemannian manifold $(M,g)$, or a probability measure $d\mu$ on a space $X$, etc.), can we effectively reconstruct $X$ given only the information of some auxiliary measurements? What if we can only perform finitely many measurements? What if the measurements are noisy? Can we reconstruct $X$ at least approximately? Can we measure in some precise way, how close our approximate reconstruction is to the unknown object $X$? And so on, and so forth.
Anyway, this post is about a cute observation, which I was reminded of while reading a paper on the inverse Gel'fand problem. Let $M$ be a compact manifold with smooth boundary $\partial M$. Then with no additional data required, we have a Banach space $L^\infty(\partial M)$ consisting of the essentially bounded measureable functions on the boundary. Since it is a Banach space, it comes with a complete metric $d_\infty(f,g) := \|f-g\|_{L^\infty(\partial M)}$.
Now, suppose that $g$ is a Riemannian metric on $M$. Then we have the Riemannian distance function $d_g(x,y)$ which is defined to be the infimum of arclengths of all smooth paths connecting $x$ and $y$. For any $x \in M$, we obtain a function $r_x \in L^\infty(\partial M)$ defined by
\[ r_x(z) = d_g(x,z), \forall z \in \partial M. \]
This gives a map $\phi_g: M \to L^\infty(\partial M)$, defined by $x \mapsto r_x$.
Theorem. Suppose that for any two distinct $x,y \in M$, there is a unique length-minimizing geodesic connecting $x$ and $y$. Then $\phi_g: M \to L^\infty(\partial M)$ is an isometric embedding, i.e. $d_g(x,y) = d_\infty(r_x, r_y)$ for all $x,y \in M$.
Proof. Let $x,y$ be distinct and let $\gamma$ be the unique geodesic from $x$ to $y$. For any point $z$ on the boundary, we have
\[ |d_g(x,z) - d_g(y,z)| \leq d_g(x,y). \]
which is the triangle inequality. Now let $\gamma$ be the unique geodesic from $x$ to $y$, and extend $\gamma$ until it hits some boundary point $z_\ast$. Then since $x,y,z_\ast$ all lie on a length-minimizing geodesic, we have
\[ d_g(x,z_\ast) - d_g(y,z_\ast) = d_g(x,y). \]
Therefore, the bound above is always saturated, and we find
\[ \sup_{z \in \partial M} |d_g(x,z) - d_g(y,z)| = d_g(x,y). \]
But the expression on the left is nothing but the $L^\infty(\partial M)$-norm of $r_x-r_y$, so the theorem is proved.
Anyway, this post is about a cute observation, which I was reminded of while reading a paper on the inverse Gel'fand problem. Let $M$ be a compact manifold with smooth boundary $\partial M$. Then with no additional data required, we have a Banach space $L^\infty(\partial M)$ consisting of the essentially bounded measureable functions on the boundary. Since it is a Banach space, it comes with a complete metric $d_\infty(f,g) := \|f-g\|_{L^\infty(\partial M)}$.
Now, suppose that $g$ is a Riemannian metric on $M$. Then we have the Riemannian distance function $d_g(x,y)$ which is defined to be the infimum of arclengths of all smooth paths connecting $x$ and $y$. For any $x \in M$, we obtain a function $r_x \in L^\infty(\partial M)$ defined by
\[ r_x(z) = d_g(x,z), \forall z \in \partial M. \]
This gives a map $\phi_g: M \to L^\infty(\partial M)$, defined by $x \mapsto r_x$.
Theorem. Suppose that for any two distinct $x,y \in M$, there is a unique length-minimizing geodesic connecting $x$ and $y$. Then $\phi_g: M \to L^\infty(\partial M)$ is an isometric embedding, i.e. $d_g(x,y) = d_\infty(r_x, r_y)$ for all $x,y \in M$.
Proof. Let $x,y$ be distinct and let $\gamma$ be the unique geodesic from $x$ to $y$. For any point $z$ on the boundary, we have
\[ |d_g(x,z) - d_g(y,z)| \leq d_g(x,y). \]
which is the triangle inequality. Now let $\gamma$ be the unique geodesic from $x$ to $y$, and extend $\gamma$ until it hits some boundary point $z_\ast$. Then since $x,y,z_\ast$ all lie on a length-minimizing geodesic, we have
\[ d_g(x,z_\ast) - d_g(y,z_\ast) = d_g(x,y). \]
Therefore, the bound above is always saturated, and we find
\[ \sup_{z \in \partial M} |d_g(x,z) - d_g(y,z)| = d_g(x,y). \]
But the expression on the left is nothing but the $L^\infty(\partial M)$-norm of $r_x-r_y$, so the theorem is proved.
Monday, August 31, 2015
Hamilton-Jacobi equation and Riemannian distance
Consider the cotangent bundle $T^\ast X$ as a symplectic manifold with canonical symplectic form $\omega$. Consider the Hamilton-Jacobi equation
\[ \frac{\partial S}{\partial t} + H(x, \nabla S) = 0, \]
for the classical Hamilton function $S(x,t)$. Setting $x=x(t), p(t) = (\nabla S)(x(t), t)$ one sees immediately from the method of characteristics that this PDE is solved by the classical action
\[ S(x,t) = \int_0^t (p \dot{x} - H) ds, \]
where the integral is taken over the solution $(x(s),p(s))$ of Hamilton's equations with $x(0)=x_0$ and $x(t) = x$. The choice of basepoint $x_0$ involves an overall additive constant of $S$, and really this solution is only valid in some neighbourhood $U$ of $x_0$. (Reason: $S$ is in general multivalued, as the differential "$dS$" is closed but not necessarily exact.)
Now consider the case where $X$ is Riemannian, with Hamiltonian $H(x,p) = \frac{1}{2} |p|^2$. The solutions to Hamilton's equations are affinely parametrized geodesics, and by a simple Legendre transform we have
\[ S(x, t) = \frac{1}{2} \int_0^t |\dot x|^2 ds \]
where the integral is along the affine geodesic with $x(0) = x_0$ and $x(t) = x$. Since $x(s)$ is a geodesic, $|\dot x(s)|$ is a constant (in $s$) and therefore
\[ S(x, t) = \frac{t}{2} |\dot x(0)|^2. \]
Now consider the path $\gamma(s) = x($|\dot x(0)|^{-1}$s)$. This is an affine geodesic with $\gamma(0) = x_0$, $\gamma(|\dot x(0)|t) = x$ and $|\dot \gamma| = 1$. Therefore, the Riemannian distance between $x_0$ and $x$ (provided $x$ is sufficiently close to $x_0$) is
\[ d(x_0, x) = |\dot x(0)| t. \]
Combining this with the previous calculation, we see that
\[ S(x, t) = \frac{1}{2t} d(x_0, x)^2. \]
Now insert this back into the Hamilton-Jacobi equation above. With a bit of rearranging, we have the following.
Theorem. Let $x_0$ denote a fixed basepoint of $X$. Then for all $x$ in a sufficiently small neighborhood $U$ of $x_0$, the Riemannian distance function satisfies the Eikonal equation
\[ |\nabla_x d(x_0, x)|^2 = 1. \]
Now, for convenience set $r(x) = d(x_0, x)$. Then $|\nabla r|^2 = 1$, from which we obtain (by differentiating twice and contracting)
\[ g^{ij} g^{kl}\left(\nabla_{lki} r \nabla_j r + \nabla_{ki}r \nabla_{lj} r\right) = 0.\]
Quick calculation shows that
\[ \nabla_{lki} r = \nabla_{ilk} r - \left.R_{li}\right.^{b}_k \nabla_b r \]
Therefore, tracing over $l$ and $k$ we obtain
\[ g^{lk} \nabla_{lki} r = \nabla_i ( \Delta r) + Rc(\nabla r, -) \]
Plugging this back into the equation derived above, we have
\[ \nabla r \cdot \nabla(\Delta r) + Rc(\nabla r, \nabla r) + |Hr|^2 = 0, \]
where $Hr$ denotes the Hessian of $r$ regarded as a 2-tensor. Now, using $r$ as a local coordinate, it is easy to see that $\partial_r = \nabla r$ (as vector fields). So we can rewrite this identity as
\[ \partial_r (\Delta r) + Rc(\partial_r, \partial_r) + |Hr|^2 = 0. \]
Now, we can get a nice result out of this. First, note that the Hessian $Hr$ always has at least one eigenvalue equal to zero, because the Eikonal equation implies that $Hr(\partial_r, -)=0$. Let $\lambda_2, \dots, \lambda_n$ denote the non-zero eigenvalues of $Hr$. We have
\[ |Hr|^2 = \lambda_2^2 + \dots + \lambda_n^2, \]
while on the other hand
\[ |\Delta r|^2 = (\lambda_2 + \dots + \lambda_n)^2 \]
By Cauchy-Schwarz, we have
\[ |\Delta r|^2 \leq (n-1)|Hr|^2 \]
Proposition. Suppose that the Ricci curvature of $X$ satisfies $Rc \geq (n-1)\kappa$, and let $u = (n-1)(\Delta r)^{-1}$. Then
\[ u' \geq 1 + \kappa u^2. \]
Proof. From preceding formulas, $|Hr|^2$ can be expressed in terms of the Ricci curvature and the radial derivative of $\Delta r$. On the other hand, $|\Delta|^2$ is bounded above by $(n-1) |Hr|^2$. The claimed inequality then follows from simple rearrangement.
Now, the amazing thing is that this deceptively simple inequality is the main ingredient of the Bishop-Gromov comparison theorem. The Bishop-Gromov comparison theorem, in turn, is the main ingredient of the proof of Gromov(-Cheeger) precompactness. I hope to discuss these topics in a future post.
\[ \frac{\partial S}{\partial t} + H(x, \nabla S) = 0, \]
for the classical Hamilton function $S(x,t)$. Setting $x=x(t), p(t) = (\nabla S)(x(t), t)$ one sees immediately from the method of characteristics that this PDE is solved by the classical action
\[ S(x,t) = \int_0^t (p \dot{x} - H) ds, \]
where the integral is taken over the solution $(x(s),p(s))$ of Hamilton's equations with $x(0)=x_0$ and $x(t) = x$. The choice of basepoint $x_0$ involves an overall additive constant of $S$, and really this solution is only valid in some neighbourhood $U$ of $x_0$. (Reason: $S$ is in general multivalued, as the differential "$dS$" is closed but not necessarily exact.)
Now consider the case where $X$ is Riemannian, with Hamiltonian $H(x,p) = \frac{1}{2} |p|^2$. The solutions to Hamilton's equations are affinely parametrized geodesics, and by a simple Legendre transform we have
\[ S(x, t) = \frac{1}{2} \int_0^t |\dot x|^2 ds \]
where the integral is along the affine geodesic with $x(0) = x_0$ and $x(t) = x$. Since $x(s)$ is a geodesic, $|\dot x(s)|$ is a constant (in $s$) and therefore
\[ S(x, t) = \frac{t}{2} |\dot x(0)|^2. \]
Now consider the path $\gamma(s) = x($|\dot x(0)|^{-1}$s)$. This is an affine geodesic with $\gamma(0) = x_0$, $\gamma(|\dot x(0)|t) = x$ and $|\dot \gamma| = 1$. Therefore, the Riemannian distance between $x_0$ and $x$ (provided $x$ is sufficiently close to $x_0$) is
\[ d(x_0, x) = |\dot x(0)| t. \]
Combining this with the previous calculation, we see that
\[ S(x, t) = \frac{1}{2t} d(x_0, x)^2. \]
Now insert this back into the Hamilton-Jacobi equation above. With a bit of rearranging, we have the following.
Theorem. Let $x_0$ denote a fixed basepoint of $X$. Then for all $x$ in a sufficiently small neighborhood $U$ of $x_0$, the Riemannian distance function satisfies the Eikonal equation
\[ |\nabla_x d(x_0, x)|^2 = 1. \]
Now, for convenience set $r(x) = d(x_0, x)$. Then $|\nabla r|^2 = 1$, from which we obtain (by differentiating twice and contracting)
\[ g^{ij} g^{kl}\left(\nabla_{lki} r \nabla_j r + \nabla_{ki}r \nabla_{lj} r\right) = 0.\]
Quick calculation shows that
\[ \nabla_{lki} r = \nabla_{ilk} r - \left.R_{li}\right.^{b}_k \nabla_b r \]
Therefore, tracing over $l$ and $k$ we obtain
\[ g^{lk} \nabla_{lki} r = \nabla_i ( \Delta r) + Rc(\nabla r, -) \]
Plugging this back into the equation derived above, we have
\[ \nabla r \cdot \nabla(\Delta r) + Rc(\nabla r, \nabla r) + |Hr|^2 = 0, \]
where $Hr$ denotes the Hessian of $r$ regarded as a 2-tensor. Now, using $r$ as a local coordinate, it is easy to see that $\partial_r = \nabla r$ (as vector fields). So we can rewrite this identity as
\[ \partial_r (\Delta r) + Rc(\partial_r, \partial_r) + |Hr|^2 = 0. \]
Now, we can get a nice result out of this. First, note that the Hessian $Hr$ always has at least one eigenvalue equal to zero, because the Eikonal equation implies that $Hr(\partial_r, -)=0$. Let $\lambda_2, \dots, \lambda_n$ denote the non-zero eigenvalues of $Hr$. We have
\[ |Hr|^2 = \lambda_2^2 + \dots + \lambda_n^2, \]
while on the other hand
\[ |\Delta r|^2 = (\lambda_2 + \dots + \lambda_n)^2 \]
By Cauchy-Schwarz, we have
\[ |\Delta r|^2 \leq (n-1)|Hr|^2 \]
Proposition. Suppose that the Ricci curvature of $X$ satisfies $Rc \geq (n-1)\kappa$, and let $u = (n-1)(\Delta r)^{-1}$. Then
\[ u' \geq 1 + \kappa u^2. \]
Proof. From preceding formulas, $|Hr|^2$ can be expressed in terms of the Ricci curvature and the radial derivative of $\Delta r$. On the other hand, $|\Delta|^2$ is bounded above by $(n-1) |Hr|^2$. The claimed inequality then follows from simple rearrangement.
Now, the amazing thing is that this deceptively simple inequality is the main ingredient of the Bishop-Gromov comparison theorem. The Bishop-Gromov comparison theorem, in turn, is the main ingredient of the proof of Gromov(-Cheeger) precompactness. I hope to discuss these topics in a future post.
Tuesday, August 18, 2015
The Classical Partition Function
Let $(M, \omega)$ be a symplectic manifold of dimension $2n$, and let $H: M \to \mathbf{R}$ be a classical Hamiltonian. The symplectic form $\omega$ allows us to define a measure on $M$, given by integration against the top form $\omega^n / n!$. We will denote this measure by $d\mu$.
We imagine that $(M, \omega, H)$ represents some classical mechanical system. We suppose that the dynamics of this dynamical system are very complicated, e.g. some system of $10^{23}$ particles. The system is so complicated that not only can we not solve the equations of motion exactly, and even if we could, their solutions might be so complicated that we can't expect to learn very much from them.
So instead, we ask statistical questions. Imagine that we cannot measure the state of the system exactly (e.g. particles in a box), so we try to guess a probability distribution $\rho(x,p,t)$ on $M$ indicating that at time $t$ the system has probability $\rho(x,p,t) d\mu$ of being in the state $(x,p)$. Obviously, $\rho$ should satisfy the constraint $\int_M \rho d\mu = 1$.
How does $\rho$ evolve in time? We know that the system obeys Hamilton's equations,
\[ (\dot x, \dot p) = X_H = (\partial H / \partial p, -\partial H / \partial x) \]
in local Darboux coordinates. Therefore, a particle located at $(x,p)$ in phase space at time $t$ will be located at $(x,p)+X_H dt$ in phase space at time $t+dt$. Therefore, the probability that a particle is at point $(x,p)$ at time $t+dt$, should be equal to the probability that the particle is at point $(x,p)-X_H dt$ at time $t$. Therefore, we have
\[ \frac{\partial \rho}{\partial t} = \frac{\partial H}{\partial x} \frac{\partial \rho}{\partial p} - \frac{\partial H}{\partial p} \frac{\partial \rho}{\partial x} = \{H, \rho\} \]
Given a probability distribution $\rho$, the entropy is defined to be
\[ S[\rho] = -\int_M \rho \log \rho d\mu. \]
(A version of) the second law of thermodynamics. For a given average energy $U$, the system assumes a distribution of maximal possible entropy at thermodynamic equilibrium.
The goal now, is to determine what distribution $\rho$ will maximize the entropy, subject to the constraints (for fixed $U$)
\begin{align*} \int_M H \rho d\mu &= U \\\
\int_M \rho d\mu &= 1 \end{align*}
Setting aside technical issues of convergence, etc., this variational problem is easily solved using the method of Lagrange multipliers. Introducing parameters $\lambda_1, \lambda_2$, we consider the modified functional
\[ S[\rho, \lambda_1, \lambda_2, U] = \int_M\left(-\rho \log \rho +\lambda_1\rho +\lambda_2(H\rho)\right)d\mu -\lambda_1-\lambda_2 U. \]
Note that $\partial S / \partial U = -\lambda_2$, and this is conventionally identified with (minus) the inverse temperature.
Taking the variation with respect to $\rho$, we find
\[ 0= \frac{\delta S}{\delta \rho} = -\log \rho-1+\lambda_1+H\lambda_2\]
Therefore, $rho$ is proportional to $e^{-\beta H}$ where we have set $\beta=-\lambda_2$. Define the partition function $Z$ to be
\[ Z = \int_M e^{-\beta H} d\mu. \]
We therefore have proved (formally and heuristically only!):
Theorem. The probability distribution $\rho$ assumed by the system at thermodynamic equilibrium is given by
\[ \rho = \frac{e^{-\beta H}}{Z} \]
where $\beta > 0$ is a real parameter, called the inverse temperature.
Corollary. At thermodynamic equilibrium, the average energy is given by
\[ U = -\frac{\partial \log Z}{\partial \beta} , \]
and the entropy is given by
\[ S = \beta U + \log Z.\]
We imagine that $(M, \omega, H)$ represents some classical mechanical system. We suppose that the dynamics of this dynamical system are very complicated, e.g. some system of $10^{23}$ particles. The system is so complicated that not only can we not solve the equations of motion exactly, and even if we could, their solutions might be so complicated that we can't expect to learn very much from them.
So instead, we ask statistical questions. Imagine that we cannot measure the state of the system exactly (e.g. particles in a box), so we try to guess a probability distribution $\rho(x,p,t)$ on $M$ indicating that at time $t$ the system has probability $\rho(x,p,t) d\mu$ of being in the state $(x,p)$. Obviously, $\rho$ should satisfy the constraint $\int_M \rho d\mu = 1$.
How does $\rho$ evolve in time? We know that the system obeys Hamilton's equations,
\[ (\dot x, \dot p) = X_H = (\partial H / \partial p, -\partial H / \partial x) \]
in local Darboux coordinates. Therefore, a particle located at $(x,p)$ in phase space at time $t$ will be located at $(x,p)+X_H dt$ in phase space at time $t+dt$. Therefore, the probability that a particle is at point $(x,p)$ at time $t+dt$, should be equal to the probability that the particle is at point $(x,p)-X_H dt$ at time $t$. Therefore, we have
\[ \frac{\partial \rho}{\partial t} = \frac{\partial H}{\partial x} \frac{\partial \rho}{\partial p} - \frac{\partial H}{\partial p} \frac{\partial \rho}{\partial x} = \{H, \rho\} \]
Given a probability distribution $\rho$, the entropy is defined to be
\[ S[\rho] = -\int_M \rho \log \rho d\mu. \]
(A version of) the second law of thermodynamics. For a given average energy $U$, the system assumes a distribution of maximal possible entropy at thermodynamic equilibrium.
The goal now, is to determine what distribution $\rho$ will maximize the entropy, subject to the constraints (for fixed $U$)
\begin{align*} \int_M H \rho d\mu &= U \\\
\int_M \rho d\mu &= 1 \end{align*}
Setting aside technical issues of convergence, etc., this variational problem is easily solved using the method of Lagrange multipliers. Introducing parameters $\lambda_1, \lambda_2$, we consider the modified functional
\[ S[\rho, \lambda_1, \lambda_2, U] = \int_M\left(-\rho \log \rho +\lambda_1\rho +\lambda_2(H\rho)\right)d\mu -\lambda_1-\lambda_2 U. \]
Note that $\partial S / \partial U = -\lambda_2$, and this is conventionally identified with (minus) the inverse temperature.
Taking the variation with respect to $\rho$, we find
\[ 0= \frac{\delta S}{\delta \rho} = -\log \rho-1+\lambda_1+H\lambda_2\]
Therefore, $rho$ is proportional to $e^{-\beta H}$ where we have set $\beta=-\lambda_2$. Define the partition function $Z$ to be
\[ Z = \int_M e^{-\beta H} d\mu. \]
We therefore have proved (formally and heuristically only!):
Theorem. The probability distribution $\rho$ assumed by the system at thermodynamic equilibrium is given by
\[ \rho = \frac{e^{-\beta H}}{Z} \]
where $\beta > 0$ is a real parameter, called the inverse temperature.
Corollary. At thermodynamic equilibrium, the average energy is given by
\[ U = -\frac{\partial \log Z}{\partial \beta} , \]
and the entropy is given by
\[ S = \beta U + \log Z.\]
Thursday, January 29, 2015
What is generalized geometry?
The following are my notes for a short introductory talk. References below are not intended to be comprehensive!
Math references:
Physics references:
What is geometry?
Before trying to define generalized geometry, we should first decide what we mean by ordinary geometry. Of course, this question doesn't have a unique answer, so there are many ways to generalize the classical notions of manifolds and varieties. The viewpoint taken in generalized geometry is the following: the distinguishing feature of smooth manifolds is the existence of a tangent bundle
\[ TM \to M \]
which satisfies some nice axioms. The basic idea of generalized geometry is to replace the tangent bundle with some other vector bundle $L \to M$, again satisfying some nice axioms. Different generalized geometries on $M$ will correspond to different choices of bundle $L \to M$, as well as auxiliary data compatible with $L$ in some appropriate sense.
Definition. A Lie algebroid over $M$ is a smooth vector bundle $L \to M$ together with a vector bundle map $a: L \to TM$ called the anchor map and a bracket $[\cdot, \cdot]: H^0(M, L) \otimes H^0(M, L) \to H^0(M, L)$ satisfying the following axioms:
Example 1. We can take $L$ to be $TM$ with anchor map the identity.
Example 2. Let $\sigma$ be a Poisson tensor on $M$. Then we define a bracket by $[X,Y] = \sigma(X,Y)$ and an anchor by $X \mapsto \sigma(X, \cdot)$. This makes $T^\ast M$ into a Lie algebroid.
Example 3. Let $M$ be a complex manifold of and let $L \subset TM \otimes \mathbf{C}$ be the sub-bundle of vectors spanned by $\{\partial / \partial z_1, \dots, \partial / \partial z_n\}$ in local holomorphic coordinates. Then $L \to M$ is a (complex) Lie algebroid.
Courant Bracket
We'd like to try to fit the preceding examples into a common framework. Let $\mathbf{T}M = TM \oplus T^\ast M$. This bundle has a natural symmetric bilinear pairing given by
\[ \langle X \oplus \alpha, Y \oplus \beta \rangle = \frac{1}{2} \alpha(Y) + \frac{1}{2} \beta(X) \]
Note that this bilinear form is of split signature $(n,n)$. We define a bracket on sections of $\mathbf{T}M$ by
\[ [X\oplus \alpha, Y\oplus \beta] = [X,Y] \oplus \left(L_X \beta + \frac{1}{2}(d \alpha(Y))- L_Y \alpha -\frac{1}{2} d( \beta(X)) \right ) \]
Note that this bracket is not a Lie bracket. We also have an anchor map $a: \mathbf{T}M \to TM$ which is just the projection.
Let $B$ be a 2-form on $M$. Define an action of $B$ on sections of $\mathbf TM$ by
\[ X + \alpha \mapsto X + \alpha + i_X B \]
Proposition. This action preserves the Courant bracket if and only if $B$ is closed.
This shows that the diffeomorphisms of $M$ as a generalized manifold are large than the ordinary diffeomorphisms of $M$. In fact is is the semidirect product of the diffeomorphism group of $M$ with the vector space of closed 2-forms.
Dirac Structures
Definition. A Dirac structure on $M$ is an Lagrangian sub-bundle $L \subset \mathbf{T}M$ which is closed under the Courant bracket.
Theorem (Courant). A Lagrangian sub-bundle $L \subset \mathbf{T} M$ is a Dirac structure if and only if $L \to M$ is a Lie algebroid over $M$, with bracket induced by the Courant bracket and anchor given by projection.
Example 1. $TM \subset \mathbf{T}M$.
Example 2. Take $L$ to be the graph of a Poisson tensor.
Example 3. Take $L$ to be the graph of a closed 2-form.
Admissible Functions
We now let $L \to M$ be a Dirac structure on $M$.
Definition. A smooth function $f$ on $M$ is called admissible if there exists a vector field $X_f$ such that $(X_f, df)$ is a section of $L$.
The Poisson bracket is defined as follows. If $f,g$ are admissible, then define
\[ \{f, g\} = X_f g. \]
It is easy to check from the definitions that the bracket on admissible functions is well-defined (independent of choice of $X_f$) and skew-symmetric. With a little bit of calculation, we find the following.
Proposition. The vector space of admissible functions is naturally a Poisson algebra, and moreover the natural bracket satisfies the Leibniz rule.
Generalized Complex Structures
Definition. A generalized complex structure is a skew endomorphism $J$ of $\mathbf T M$ such that $J^2 = -1$ and such that the $+i$-eigenbundle is involutive under the Courant bracket.
Equivalently: A generalized complex structure is a (complex) Dirac structure $L \subset \mathbf TM$ satisfying the condition $L \cap \overline L = 0$.
Example 1. Let $J$ be an ordinary complex structure on $M$. Then the endomorphism
\[ \begin{bmatrix} -J & 0 \\ 0 & J^\ast \end{bmatrix} \]
defines a generalized complex structure on $M$.
Example 2. Let $\omega$ be a symplectic form on $M$. Then the endomorphism
\[ \begin{bmatrix} 0 & -\omega^{-1} \\ \omega & 0 \end{bmatrix} \]
defines a generalized complex structure on $M$.
Thus, generalized geometry gives a common framework for both complex geometry and symplectic geometry. Such a connection is exactly what is conjectured by mirror symmetry.
Example 3. Let $J$ be a complex structure on $M$ and let $\sigma$ be a holomorphic Poisson tensor. Consider the subbundle $L \subset \mathbf TM$ defined as the span of
\[ \frac{\partial}{\partial \bar z_1}, \dots, \frac{\partial}{\partial \bar z_n}, dz_1 - \sigma(dz_1), \dots, dz_n - \sigma(dz_n) \]
Then $L$ defines a generalized complex structure on $M$.
The last example shows that deformations of $M$ as a generalized complex manifold contain non-commutative deformations of the structure sheaf. We also have the following theorem, which shows that there is an intimate relation between generalized complex geometry and holomorphic Poisson geometry.
Theorem (Bailey). Near any point of a generalized complex manifold, $M$ is locally isomorphic to the product of a holomorphic Poisson manifold with a symplectic manifold.
Generalized Kähler Manifolds
Let $(g, J, \omega)$ be a Kähler triple. The Kähler property requires that
\[ \omega = g J. \]
Let $I_1$ denote the generalized complex structure induced by $J$, and let $I_1$ denote the generalized complex structure induced by the symplectic form $\omega$. We have
\[ I_1 I_2 = \begin{bmatrix} - J & 0 \\ 0 & J^\ast \end{bmatrix} \begin{bmatrix} 0 & -\omega^{-1} \\ \omega & 0 \end{bmatrix} = \begin{bmatrix} 0 & g^{-1} \\ g & 0 \end{bmatrix} = I_2 I_1 \]
Definition. A generalized Kähler manifold is a manifold with two commuting generalized complex structure $I_1, I_2$ such that the bilinear pairing $(I_1 I_2 u, v)$ is positive definite.
Theorem (Gualtieri). A generalized Kähler structure on $M$ induces a Riemannian metric $g$, two integrable almost complex structures $J_\pm$ Hermitian with respect to $g$, and two affine connections $\nabla_\pm$ with skew-torsion $\pm H$ which preserve the metric and complex structure $J_\pm$. Conversely, these data determine a generalized Kähler structure which is unique up to a B-field transformation.
Thus the notion of generalized Kähler manifold recovers the bihermitian geometry investigated by physicists in the context of susy non-linear $\sigma$-models.
Generalized Calabi-Yau Manifolds
Definition. A generalized Calabi-Yau manifold is a manifold $M$ together with a complex-valued differential form $\phi$, which is either purely even or purely odd, which is a pure spinor for the action of $Cl(\mathbf TM)$ and satisfies the non-degeneracy condition $(\phi, \bar \phi) \neq 0$.
Note that (by definition) $\phi$ is pure if its annihilator is a maximal isotropic subspace. Let $L \subset \mathbf TM$ be its annihilator. Then it is not hard to see that $L$ defines a generalized complex structure on $M$, so indeed a generalized Calabi-Yau manifold is in particular a generalized complex manifold.
Example. If $M$ is a complex manifold with a nowhere vanishing holomorphic $(n,0)$ form, then it is generalized Calabi-Yau.
Example. If $M$ is symplectic with symplectic form $\omega$, then $\phi = \exp(i\omega)$ gives $M$ the structure of a generalized Calabi-Yau manifold.
If $(M, \phi)$ is generalized Calabi-Yau, then so is $(M, \exp(B) \phi)$ for any closed real 2-form $B$. In the symplectic case, we obtain
\[ \phi = \exp(B+i\omega) \]
This explains the appearance of the $B$-field (or "complexified Kähler form") in discussions of mirror symmetry.
Math references:
- Courant, Dirac manifolds
- Hitchin, Generalized Calabi-Yau manifolds
- Gualtieri, Generalized complex geometry
- Cavalcanti, New aspects of the ddc-lemma
- Cavalcanti and Gualtieri, Generalized complex geometry and T-duality
- Bailey and Gualtieri, Local analytic geometry of generalized complex structures
Physics references:
- Dijkgraaf, Gukov, Neitzke, Vafa, Topological M-theory as unification of form theories of gravity
- Grana, Flux compactifications in string theory: a comprehensive review
What is geometry?
Before trying to define generalized geometry, we should first decide what we mean by ordinary geometry. Of course, this question doesn't have a unique answer, so there are many ways to generalize the classical notions of manifolds and varieties. The viewpoint taken in generalized geometry is the following: the distinguishing feature of smooth manifolds is the existence of a tangent bundle
\[ TM \to M \]
which satisfies some nice axioms. The basic idea of generalized geometry is to replace the tangent bundle with some other vector bundle $L \to M$, again satisfying some nice axioms. Different generalized geometries on $M$ will correspond to different choices of bundle $L \to M$, as well as auxiliary data compatible with $L$ in some appropriate sense.
Definition. A Lie algebroid over $M$ is a smooth vector bundle $L \to M$ together with a vector bundle map $a: L \to TM$ called the anchor map and a bracket $[\cdot, \cdot]: H^0(M, L) \otimes H^0(M, L) \to H^0(M, L)$ satisfying the following axioms:
- $[\cdot,\cdot]$ is a Lie bracket on $H^0(M, L)$
- $[X, fY] = f[X,Y] + a(X)f \cdot Y$ for $X,Y \in H^0(M,L)$ and $f \in H^0(M, \mathcal{O}_M)$
Example 1. We can take $L$ to be $TM$ with anchor map the identity.
Example 2. Let $\sigma$ be a Poisson tensor on $M$. Then we define a bracket by $[X,Y] = \sigma(X,Y)$ and an anchor by $X \mapsto \sigma(X, \cdot)$. This makes $T^\ast M$ into a Lie algebroid.
Example 3. Let $M$ be a complex manifold of and let $L \subset TM \otimes \mathbf{C}$ be the sub-bundle of vectors spanned by $\{\partial / \partial z_1, \dots, \partial / \partial z_n\}$ in local holomorphic coordinates. Then $L \to M$ is a (complex) Lie algebroid.
Courant Bracket
We'd like to try to fit the preceding examples into a common framework. Let $\mathbf{T}M = TM \oplus T^\ast M$. This bundle has a natural symmetric bilinear pairing given by
\[ \langle X \oplus \alpha, Y \oplus \beta \rangle = \frac{1}{2} \alpha(Y) + \frac{1}{2} \beta(X) \]
Note that this bilinear form is of split signature $(n,n)$. We define a bracket on sections of $\mathbf{T}M$ by
\[ [X\oplus \alpha, Y\oplus \beta] = [X,Y] \oplus \left(L_X \beta + \frac{1}{2}(d \alpha(Y))- L_Y \alpha -\frac{1}{2} d( \beta(X)) \right ) \]
Note that this bracket is not a Lie bracket. We also have an anchor map $a: \mathbf{T}M \to TM$ which is just the projection.
Let $B$ be a 2-form on $M$. Define an action of $B$ on sections of $\mathbf TM$ by
\[ X + \alpha \mapsto X + \alpha + i_X B \]
Proposition. This action preserves the Courant bracket if and only if $B$ is closed.
This shows that the diffeomorphisms of $M$ as a generalized manifold are large than the ordinary diffeomorphisms of $M$. In fact is is the semidirect product of the diffeomorphism group of $M$ with the vector space of closed 2-forms.
Dirac Structures
Definition. A Dirac structure on $M$ is an Lagrangian sub-bundle $L \subset \mathbf{T}M$ which is closed under the Courant bracket.
Theorem (Courant). A Lagrangian sub-bundle $L \subset \mathbf{T} M$ is a Dirac structure if and only if $L \to M$ is a Lie algebroid over $M$, with bracket induced by the Courant bracket and anchor given by projection.
Example 1. $TM \subset \mathbf{T}M$.
Example 2. Take $L$ to be the graph of a Poisson tensor.
Example 3. Take $L$ to be the graph of a closed 2-form.
Admissible Functions
We now let $L \to M$ be a Dirac structure on $M$.
Definition. A smooth function $f$ on $M$ is called admissible if there exists a vector field $X_f$ such that $(X_f, df)$ is a section of $L$.
The Poisson bracket is defined as follows. If $f,g$ are admissible, then define
\[ \{f, g\} = X_f g. \]
It is easy to check from the definitions that the bracket on admissible functions is well-defined (independent of choice of $X_f$) and skew-symmetric. With a little bit of calculation, we find the following.
Proposition. The vector space of admissible functions is naturally a Poisson algebra, and moreover the natural bracket satisfies the Leibniz rule.
Generalized Complex Structures
Definition. A generalized complex structure is a skew endomorphism $J$ of $\mathbf T M$ such that $J^2 = -1$ and such that the $+i$-eigenbundle is involutive under the Courant bracket.
Equivalently: A generalized complex structure is a (complex) Dirac structure $L \subset \mathbf TM$ satisfying the condition $L \cap \overline L = 0$.
Example 1. Let $J$ be an ordinary complex structure on $M$. Then the endomorphism
\[ \begin{bmatrix} -J & 0 \\ 0 & J^\ast \end{bmatrix} \]
defines a generalized complex structure on $M$.
Example 2. Let $\omega$ be a symplectic form on $M$. Then the endomorphism
\[ \begin{bmatrix} 0 & -\omega^{-1} \\ \omega & 0 \end{bmatrix} \]
defines a generalized complex structure on $M$.
Thus, generalized geometry gives a common framework for both complex geometry and symplectic geometry. Such a connection is exactly what is conjectured by mirror symmetry.
Example 3. Let $J$ be a complex structure on $M$ and let $\sigma$ be a holomorphic Poisson tensor. Consider the subbundle $L \subset \mathbf TM$ defined as the span of
\[ \frac{\partial}{\partial \bar z_1}, \dots, \frac{\partial}{\partial \bar z_n}, dz_1 - \sigma(dz_1), \dots, dz_n - \sigma(dz_n) \]
Then $L$ defines a generalized complex structure on $M$.
The last example shows that deformations of $M$ as a generalized complex manifold contain non-commutative deformations of the structure sheaf. We also have the following theorem, which shows that there is an intimate relation between generalized complex geometry and holomorphic Poisson geometry.
Theorem (Bailey). Near any point of a generalized complex manifold, $M$ is locally isomorphic to the product of a holomorphic Poisson manifold with a symplectic manifold.
Generalized Kähler Manifolds
Let $(g, J, \omega)$ be a Kähler triple. The Kähler property requires that
\[ \omega = g J. \]
Let $I_1$ denote the generalized complex structure induced by $J$, and let $I_1$ denote the generalized complex structure induced by the symplectic form $\omega$. We have
\[ I_1 I_2 = \begin{bmatrix} - J & 0 \\ 0 & J^\ast \end{bmatrix} \begin{bmatrix} 0 & -\omega^{-1} \\ \omega & 0 \end{bmatrix} = \begin{bmatrix} 0 & g^{-1} \\ g & 0 \end{bmatrix} = I_2 I_1 \]
Definition. A generalized Kähler manifold is a manifold with two commuting generalized complex structure $I_1, I_2$ such that the bilinear pairing $(I_1 I_2 u, v)$ is positive definite.
Theorem (Gualtieri). A generalized Kähler structure on $M$ induces a Riemannian metric $g$, two integrable almost complex structures $J_\pm$ Hermitian with respect to $g$, and two affine connections $\nabla_\pm$ with skew-torsion $\pm H$ which preserve the metric and complex structure $J_\pm$. Conversely, these data determine a generalized Kähler structure which is unique up to a B-field transformation.
Thus the notion of generalized Kähler manifold recovers the bihermitian geometry investigated by physicists in the context of susy non-linear $\sigma$-models.
Generalized Calabi-Yau Manifolds
Definition. A generalized Calabi-Yau manifold is a manifold $M$ together with a complex-valued differential form $\phi$, which is either purely even or purely odd, which is a pure spinor for the action of $Cl(\mathbf TM)$ and satisfies the non-degeneracy condition $(\phi, \bar \phi) \neq 0$.
Note that (by definition) $\phi$ is pure if its annihilator is a maximal isotropic subspace. Let $L \subset \mathbf TM$ be its annihilator. Then it is not hard to see that $L$ defines a generalized complex structure on $M$, so indeed a generalized Calabi-Yau manifold is in particular a generalized complex manifold.
Example. If $M$ is a complex manifold with a nowhere vanishing holomorphic $(n,0)$ form, then it is generalized Calabi-Yau.
Example. If $M$ is symplectic with symplectic form $\omega$, then $\phi = \exp(i\omega)$ gives $M$ the structure of a generalized Calabi-Yau manifold.
If $(M, \phi)$ is generalized Calabi-Yau, then so is $(M, \exp(B) \phi)$ for any closed real 2-form $B$. In the symplectic case, we obtain
\[ \phi = \exp(B+i\omega) \]
This explains the appearance of the $B$-field (or "complexified Kähler form") in discussions of mirror symmetry.
Thursday, December 11, 2014
Virasoro Algebra from the Free Boson without Regularization
Usual physics derivations of the Virasoro algebra from the free boson in two dimensions usually use some sort of regularization procedure to compute the central charge. Following these notes I'd like to give a purely algebraic calculation of the central charge.
Let $A = \mathbf C[x_1, x_2, \dots]$ be the polynomial algebra in countably many generators. For an integer $k > 0$, define a $k$-linear operator $a_k$ on $A$ by
\[ a_k = \frac{\partial}{\partial x_k}, \ k > 0 \]
Similarly, for $k < 0$ we define $a_k$ by multiplication:
\[ a_{-k} = k x_k, k > 0 \]
For $k = 0$, we define $a_0$ to be multiplication by some fixed complex number (which by abuse of notation we also denote by $a_0$).
Lemma. We have the commutation relation $[a_m, a_n] = m \delta_{m+n}$ as linear operators on $A$.
For any monomial in the $a_k$, we define normal ordering $::$ to be the monomial obtained by reordering the terms so that the indices are increasing. (Mathematical interpretation: it is a section of the quotient map from the tensor algebra in the $a_k$ to the symmetric algebra, defined by lexicographic order.) For example,
\[ :a_j a_k:\ = \left\{ \begin{array}{rr} a_j a_k, & j \leq k \\ a_k a_j, & j > k \end{array} \right. \]
Next we formally define a set of operators $L_k$ by
\[ L_k = \frac{1}{2}\sum_j :a_j a_{k-j}: \]
Proposition. The $L_k$ are well-defined as linear operators on $A$.
Proof. For sufficiently large $|j|$, at least one of $j$ or $k-j$ is positive, and hence $:a_j a_{k-j}:$ contains a differentiation (on the right). Since any element $f \in A$ is annihilated by all but finitely many of the differentiation operators $\partial_j$, the formal expression $L_k f$ contains only finitely many non-zero terms, and hence is well-defined.
Lemma. As operators on $A$, we have $[a_k, L_n] = k a_{k+n}$.
Theorem. As operators on $A$, we have
\[ [L_m, L_n] = (m-n) L_{m+n} + \frac{1}{12} (m^3-m) \delta_{m+n} \]
Proof. Fix $m,n$. For the sake of simplicity we will assume $m \neq n$ and $mn \neq 0$. (The other special cases can be treated by similar arguments.) By the same argument as the proof of the preceding proposition, for any fixed element $f \in A$ there exists some $N \gg 0$ such that
\[ [L_m, L_n]f = [L_m^N, L_n] f \]
where $L_m^N$ is the truncated operator
\[ L_m^N = \frac{1}{2}\sum_{|j| < N} :a_j a_{m-j}: \]
Let us compute (noting that since $m \neq 0$, $:a_j a_{m-j}: = a_j a_{m-j}$)
\begin{align}
[L_m^N, L_n] &= \frac{1}{2} \sum_{|j| < N} [a_j a_{m-j}, L_n] \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) a_j a_{m+n-j} + \frac{1}{2} \sum_{|j| < N}j a_{n+j} a_{m-j}
\end{align}
Denote the two sums above by $S_1$ and $S_2$. It is clear that these should be related to the operator $L_{m+n}$, but to see the exact relation we will have to normal order the terms. Let's start with $S_1$. Note that $a_j a_{m+n-j}$ is already normal ordered, unless $j > m+n-j$. Hence
\begin{align}
S_1 &= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{1}{2} \sum_{m+n\lt2j\lt2N} (m-j) [a_j, a_{m+n-j}] \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{\delta_{m+n}}{2} \sum_{m+n\lt2j\lt2N} j(m-j) \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{\delta_{m+n}}{2} \sum_{0\lt j\lt N} j(m-j)
\end{align}
Similarly, we normal order the terms in $S_2$:
\begin{align}
S_2 &= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{1}{2} \sum_{m-n\lt2j\lt2N}j [a_{m-j}, a_{n+j}] \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j) \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j) \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j)
\end{align}
Hence we have
\[ S_1 + S_2 = \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: + \frac{\delta_{m+n}}{2} \sum_{0\lt j\leq m} j(m-j) \]
Now that everything is normal ordered, we can take the limit $N \to \infty$ without fear. After a simple cancellation, and explicitly summing the last term using well-known formulas for sums of powers of integers, we obtain:
\[ [L_m, L_n] = (m-n) L_{m+n} + \frac{1}{12} (m^3-m) \delta_{m+n} \]
For the usual conventions of the Virasoro algebra, this shows that this representation corresponds to central charge $c=1$.
Remark. If we tried to take the limit $N \to \infty$ in each of the terms $S_1, S_2$ separately, before taking their sum, we would obtain a formal infinite constant $\sum_{j} j(m-j)$. In any physics textbook, the author will simply zeta-regularize this sum to obtain a fininte result. However, the above calculation shows that this is not necessary. By taking care that each term in the expression $S_1+S_2$ was normal-ordered, before taking the limit, we obtain only finite constants with no need to regularize. Zeta regularization certainly has its uses (for example in rigorous definitions of functional determinants), but as the above calculation shows, it can also be an unnecessary crutch that obscures the underlying mathematical phenomena.
Remark. There is an analogous calculation which shows that one obtains a Virasoro representation from affine Lie algebras. Physically, this corresponds to the WZW model. Roughly, the generators of the affine Lie algebra behave as an infinite set of harmonic oscillators, similar to the Heisenberg algebra above. Sometime in the future I may write a sequel to this post giving the details of this calculation.
Let $A = \mathbf C[x_1, x_2, \dots]$ be the polynomial algebra in countably many generators. For an integer $k > 0$, define a $k$-linear operator $a_k$ on $A$ by
\[ a_k = \frac{\partial}{\partial x_k}, \ k > 0 \]
Similarly, for $k < 0$ we define $a_k$ by multiplication:
\[ a_{-k} = k x_k, k > 0 \]
For $k = 0$, we define $a_0$ to be multiplication by some fixed complex number (which by abuse of notation we also denote by $a_0$).
Lemma. We have the commutation relation $[a_m, a_n] = m \delta_{m+n}$ as linear operators on $A$.
For any monomial in the $a_k$, we define normal ordering $::$ to be the monomial obtained by reordering the terms so that the indices are increasing. (Mathematical interpretation: it is a section of the quotient map from the tensor algebra in the $a_k$ to the symmetric algebra, defined by lexicographic order.) For example,
\[ :a_j a_k:\ = \left\{ \begin{array}{rr} a_j a_k, & j \leq k \\ a_k a_j, & j > k \end{array} \right. \]
Next we formally define a set of operators $L_k$ by
\[ L_k = \frac{1}{2}\sum_j :a_j a_{k-j}: \]
Proposition. The $L_k$ are well-defined as linear operators on $A$.
Proof. For sufficiently large $|j|$, at least one of $j$ or $k-j$ is positive, and hence $:a_j a_{k-j}:$ contains a differentiation (on the right). Since any element $f \in A$ is annihilated by all but finitely many of the differentiation operators $\partial_j$, the formal expression $L_k f$ contains only finitely many non-zero terms, and hence is well-defined.
Lemma. As operators on $A$, we have $[a_k, L_n] = k a_{k+n}$.
Theorem. As operators on $A$, we have
\[ [L_m, L_n] = (m-n) L_{m+n} + \frac{1}{12} (m^3-m) \delta_{m+n} \]
Proof. Fix $m,n$. For the sake of simplicity we will assume $m \neq n$ and $mn \neq 0$. (The other special cases can be treated by similar arguments.) By the same argument as the proof of the preceding proposition, for any fixed element $f \in A$ there exists some $N \gg 0$ such that
\[ [L_m, L_n]f = [L_m^N, L_n] f \]
where $L_m^N$ is the truncated operator
\[ L_m^N = \frac{1}{2}\sum_{|j| < N} :a_j a_{m-j}: \]
Let us compute (noting that since $m \neq 0$, $:a_j a_{m-j}: = a_j a_{m-j}$)
\begin{align}
[L_m^N, L_n] &= \frac{1}{2} \sum_{|j| < N} [a_j a_{m-j}, L_n] \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) a_j a_{m+n-j} + \frac{1}{2} \sum_{|j| < N}j a_{n+j} a_{m-j}
\end{align}
Denote the two sums above by $S_1$ and $S_2$. It is clear that these should be related to the operator $L_{m+n}$, but to see the exact relation we will have to normal order the terms. Let's start with $S_1$. Note that $a_j a_{m+n-j}$ is already normal ordered, unless $j > m+n-j$. Hence
\begin{align}
S_1 &= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{1}{2} \sum_{m+n\lt2j\lt2N} (m-j) [a_j, a_{m+n-j}] \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{\delta_{m+n}}{2} \sum_{m+n\lt2j\lt2N} j(m-j) \\
&= \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{\delta_{m+n}}{2} \sum_{0\lt j\lt N} j(m-j)
\end{align}
Similarly, we normal order the terms in $S_2$:
\begin{align}
S_2 &= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{1}{2} \sum_{m-n\lt2j\lt2N}j [a_{m-j}, a_{n+j}] \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j) \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j) \\
&= \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: -\frac{\delta_{m+n}}{2} \sum_{m\lt j\lt N}j (m-j)
\end{align}
Hence we have
\[ S_1 + S_2 = \frac{1}{2} \sum_{|j| < N} (m-j) :a_j a_{m+n-j}: + \frac{1}{2} \sum_{|j| < N}j :a_{n+j} a_{m-j}: + \frac{\delta_{m+n}}{2} \sum_{0\lt j\leq m} j(m-j) \]
Now that everything is normal ordered, we can take the limit $N \to \infty$ without fear. After a simple cancellation, and explicitly summing the last term using well-known formulas for sums of powers of integers, we obtain:
\[ [L_m, L_n] = (m-n) L_{m+n} + \frac{1}{12} (m^3-m) \delta_{m+n} \]
For the usual conventions of the Virasoro algebra, this shows that this representation corresponds to central charge $c=1$.
Remark. If we tried to take the limit $N \to \infty$ in each of the terms $S_1, S_2$ separately, before taking their sum, we would obtain a formal infinite constant $\sum_{j} j(m-j)$. In any physics textbook, the author will simply zeta-regularize this sum to obtain a fininte result. However, the above calculation shows that this is not necessary. By taking care that each term in the expression $S_1+S_2$ was normal-ordered, before taking the limit, we obtain only finite constants with no need to regularize. Zeta regularization certainly has its uses (for example in rigorous definitions of functional determinants), but as the above calculation shows, it can also be an unnecessary crutch that obscures the underlying mathematical phenomena.
Remark. There is an analogous calculation which shows that one obtains a Virasoro representation from affine Lie algebras. Physically, this corresponds to the WZW model. Roughly, the generators of the affine Lie algebra behave as an infinite set of harmonic oscillators, similar to the Heisenberg algebra above. Sometime in the future I may write a sequel to this post giving the details of this calculation.
Wednesday, May 14, 2014
The Basic Idea of the Quantum BV Complex
$
\newcommand{\h}{\hbar}
\newcommand{\X}{\mathfrak{X}}
\newcommand{\PV}{\text{PV}}
\newcommand{\K}{{\mathbb K}}
\newcommand{\R}{{\mathbb R}}
\newcommand{\too}{\xrightarrow}
\newcommand{\im}{\text{Im }}
$
Let $M$ be an oriented $n$-dimensional manifold and $\X^\bullet(M):=\Gamma(M,\wedge^{-\bullet} TM)$, so that $\X^\bullet(M)$ is concentrated in non-positive degree. Let $\mu\in \Omega^n(M)$ a volume form on $M$. Then interior product with $\mu$ gives an isomorphism $\vee\mu: \X^k(M)\too{\cong}\Omega^{n-k}(M)$. From this, we induce a degree 1 differential $\Delta_\mu$ on $\X^\bullet(M)$ from the de Rham differential $d$ on $\Omega^\bullet(M)$, defined by $\Delta_\mu= (\vee\mu)^{-1}\circ d\circ\vee\mu$, making $\X^\bullet(M)$ into a cochain complex isomorphic to $\Omega^\bullet(M)[k]$; in particular $H^k(\X)=H^{n+k}(\Omega)$.
If $M$ is compact and connected, then $H^0(\X)=H^n(\Omega)$ is one dimensional, and upon fixing a basis to identify $H^0(\X)\too{\cong}\R$, the quotient map gives a map $\pi:C^\infty(M)\to \R$. We have the following:
Proposition Let $1\in C^\infty(M)$ be the constant function taking the value $1$. Then $[1]\in H^0(\X)$ is nontrivial and after choosing it as a basis the resulting map $\pi:C^\infty(M)\to \R$ is given by $$ f\mapsto\frac{ \int_M f\mu}{\int_M \mu}$$ Proof Let $f\in C^\infty(M)$ and suppose $f=\Delta_\mu(X)$ for $X\in\X^1(M)$. Then $$f\mu = f\vee \mu = \Delta_\mu(X)\vee\mu= d(X\vee \mu)$$ so that $f\mu$ is exact and by Stokes' theorem integrates to $0$ on $M$. Thus all $f\in \im \Delta_\mu$ integrate to zero against $\mu$ on $M$, so that the above map indeed descends to the quotient.
The above arguement also implies that were $1\in\im\Delta_\mu$ then $\mu$ would integrate to zero, contradicting that $\mu$ is a volume form, so that $[1]\in H^0(\X)$ is indeed nontrivial. The map is thus well-defined, linear, and has the appropriate action on a basis so that it is correct as claimed. $\square$
This recovers the standard definition of the expectation of an observable in the path integral picture of quantum field theory. The upshot is that having formulated the integration homologically, we can hope to extend this homological definition of expectation to situations where the integral itself is not well defined.
Consider the case $M=V$ a vector space, which in particular described the situation in coordinates on $M$. We are interested in measures of the form $\mu=e^{-S/\h}\mu_0$ where $\mu_0$ is the Lesbesgue measure on $V$, given by $\mu_0=dx^1\wedge ...\wedge dx^n$. Let $X\in \X^k(M)$, we have: \begin{align*} d(X\vee \mu) & = d(e^{-S/\h}(X\vee\mu_0)) \\ & = de^{-S/\h}\wedge (X\vee \mu_0)+e^{-S/\h}d(X\vee \mu_0)\\ & = -\frac{1}{\h}e^{-S/\h} dS\wedge (X\vee \mu_0) + e^{-S/\h} (\Delta_{\mu_0}X)\vee \mu_0 \\ & = -\frac{1}{\h}e^{-S/\h} (dS\vee X)\vee \mu_0 + (\Delta_{\mu_0}X)\vee \mu\\ & = \left(-\frac{1}{\h}dS\vee X + \Delta_{\mu_0}X \right)\vee \mu \end{align*} so that $$\Delta_\mu=\Delta_{\mu_0}-\frac{1}{\h}\iota_{dS}$$ Further, we can explicitly compute $\Delta_{\mu_0}$: for $X\in\X^k(M)$ we have that $X=\sum_I X^I \partial_I$, where the sum is over increasing $k$-tuples $I\subset\{1,...,n\}$. We have \begin{align*} d(X\vee \mu_0) & = d\left( \sum_I X^I \partial_I\vee(dx^1\wedge ...\wedge dx^n)\right)\\ & = d\left( \sum_I X^I (-1)^?dx^1\wedge ...\wedge \hat{dx^I} \wedge ...\wedge dx^n \right)\\ & = \sum_I \sum_{i\in I} \partial_{x^i} X^I (-1)^?dx^i\wedge dx^1\wedge ...\wedge \hat{ dx^I}\wedge ...\wedge dx^n\\ & = \left( \sum_i \partial_{x^i} (dx^i\vee X) \right) \vee \mu \end{align*} so that $$\Delta_{\mu_0} = \sum_i \partial_{x^i} \iota_{dx^i}$$ To be slightly more careful about the formal variable $\h$ and allow the $\h\to 0$ limit to be more clear, we refine our complex to be: $$ \X^\bullet(M)[[\h]] \quad\quad\text{equipped with}\quad\quad \h \Delta_\mu= \h\Delta_{\mu_0} -\iota_{dS}$$ Next, we restrict consideration only to polynomial observables (functions), and vector fields with polynomial coefficients, denoting the resulting complex $\PV^\bullet$. One can check that $H^0(\PV)$ is still one dimensional, so that the above proposition holds and we maintain the integral intepretation of this cohomology.
Now, we can identify $\PV^\bullet[[\h]]$ with the graded-commutative graded algebra $\K[[x^1,...,x^n,\xi_1,...,\xi_n,\h]]$ where $x^1,...,x^n,\h$ are in degree 0 and $\xi_1,...,\xi_n$ are in degree -1 as follows: $$ x^i \mapsto x^i \quad\quad \partial_i\mapsto \xi_i\quad\quad \partial_i\wedge\partial_j\mapsto \xi_i\xi_j \quad\quad \h\mapsto \h$$ Under this identification, the map $\iota_{dx^i}:PV^k(M)\to \PV^{k-1}(M)$ is identified with $\partial_{\xi_i}$. Now, we require our action function $S:M\to \R$ also be polynomial, and further, that $$S(x)=\frac{1}{2}\sum_{i,j} a_{ij}x^ix^j- b(x)$$ for $a_{ij}$ symmetric and non-degenerate and for $b\in I^3$ where $I=(x^1,...,x^n)\subset \K[[x^1,...,x^n]]$, that is, $b(x)$ a polynomial with no terms of degree less than 3. This implies that \begin{align*} \h\Delta_\mu & = \h\Delta_{\mu_0} -\iota_{dS} \\ & = \h \sum_i \partial_{x^i} \iota_{dx^i} - \sum_i (\partial_{x^i}S) \iota_{dx^i} \\ & = \h \sum_i \partial_{x^i} \iota_{dx^i} + \sum_i (\partial_{x^i}b) \iota_{dx^i} - \sum_{i,j}a_{ij}x^i\iota_{dx^j}\\ \end{align*} and under our identification this becomes $$\h\Delta_\mu = \h \sum_i \partial_{x^i}\partial_{\xi_i} + \sum_i (\partial_{x^i}b)\partial_{\xi_i} - \sum_{i,j}a_{ij}x^i\partial_{\xi_j}$$ The computation of the degree 0 cohomology of a given polynomial $f\in\K[[x_1,...,x_n]]$ under this differential is taken up in Gwilliam, Johnson-Freyd where it is shown the answer is precisely the Feynman diagram expansion for the expectation of $f$ which we expect.
If $M$ is compact and connected, then $H^0(\X)=H^n(\Omega)$ is one dimensional, and upon fixing a basis to identify $H^0(\X)\too{\cong}\R$, the quotient map gives a map $\pi:C^\infty(M)\to \R$. We have the following:
Proposition Let $1\in C^\infty(M)$ be the constant function taking the value $1$. Then $[1]\in H^0(\X)$ is nontrivial and after choosing it as a basis the resulting map $\pi:C^\infty(M)\to \R$ is given by $$ f\mapsto\frac{ \int_M f\mu}{\int_M \mu}$$ Proof Let $f\in C^\infty(M)$ and suppose $f=\Delta_\mu(X)$ for $X\in\X^1(M)$. Then $$f\mu = f\vee \mu = \Delta_\mu(X)\vee\mu= d(X\vee \mu)$$ so that $f\mu$ is exact and by Stokes' theorem integrates to $0$ on $M$. Thus all $f\in \im \Delta_\mu$ integrate to zero against $\mu$ on $M$, so that the above map indeed descends to the quotient.
The above arguement also implies that were $1\in\im\Delta_\mu$ then $\mu$ would integrate to zero, contradicting that $\mu$ is a volume form, so that $[1]\in H^0(\X)$ is indeed nontrivial. The map is thus well-defined, linear, and has the appropriate action on a basis so that it is correct as claimed. $\square$
This recovers the standard definition of the expectation of an observable in the path integral picture of quantum field theory. The upshot is that having formulated the integration homologically, we can hope to extend this homological definition of expectation to situations where the integral itself is not well defined.
Consider the case $M=V$ a vector space, which in particular described the situation in coordinates on $M$. We are interested in measures of the form $\mu=e^{-S/\h}\mu_0$ where $\mu_0$ is the Lesbesgue measure on $V$, given by $\mu_0=dx^1\wedge ...\wedge dx^n$. Let $X\in \X^k(M)$, we have: \begin{align*} d(X\vee \mu) & = d(e^{-S/\h}(X\vee\mu_0)) \\ & = de^{-S/\h}\wedge (X\vee \mu_0)+e^{-S/\h}d(X\vee \mu_0)\\ & = -\frac{1}{\h}e^{-S/\h} dS\wedge (X\vee \mu_0) + e^{-S/\h} (\Delta_{\mu_0}X)\vee \mu_0 \\ & = -\frac{1}{\h}e^{-S/\h} (dS\vee X)\vee \mu_0 + (\Delta_{\mu_0}X)\vee \mu\\ & = \left(-\frac{1}{\h}dS\vee X + \Delta_{\mu_0}X \right)\vee \mu \end{align*} so that $$\Delta_\mu=\Delta_{\mu_0}-\frac{1}{\h}\iota_{dS}$$ Further, we can explicitly compute $\Delta_{\mu_0}$: for $X\in\X^k(M)$ we have that $X=\sum_I X^I \partial_I$, where the sum is over increasing $k$-tuples $I\subset\{1,...,n\}$. We have \begin{align*} d(X\vee \mu_0) & = d\left( \sum_I X^I \partial_I\vee(dx^1\wedge ...\wedge dx^n)\right)\\ & = d\left( \sum_I X^I (-1)^?dx^1\wedge ...\wedge \hat{dx^I} \wedge ...\wedge dx^n \right)\\ & = \sum_I \sum_{i\in I} \partial_{x^i} X^I (-1)^?dx^i\wedge dx^1\wedge ...\wedge \hat{ dx^I}\wedge ...\wedge dx^n\\ & = \left( \sum_i \partial_{x^i} (dx^i\vee X) \right) \vee \mu \end{align*} so that $$\Delta_{\mu_0} = \sum_i \partial_{x^i} \iota_{dx^i}$$ To be slightly more careful about the formal variable $\h$ and allow the $\h\to 0$ limit to be more clear, we refine our complex to be: $$ \X^\bullet(M)[[\h]] \quad\quad\text{equipped with}\quad\quad \h \Delta_\mu= \h\Delta_{\mu_0} -\iota_{dS}$$ Next, we restrict consideration only to polynomial observables (functions), and vector fields with polynomial coefficients, denoting the resulting complex $\PV^\bullet$. One can check that $H^0(\PV)$ is still one dimensional, so that the above proposition holds and we maintain the integral intepretation of this cohomology.
Now, we can identify $\PV^\bullet[[\h]]$ with the graded-commutative graded algebra $\K[[x^1,...,x^n,\xi_1,...,\xi_n,\h]]$ where $x^1,...,x^n,\h$ are in degree 0 and $\xi_1,...,\xi_n$ are in degree -1 as follows: $$ x^i \mapsto x^i \quad\quad \partial_i\mapsto \xi_i\quad\quad \partial_i\wedge\partial_j\mapsto \xi_i\xi_j \quad\quad \h\mapsto \h$$ Under this identification, the map $\iota_{dx^i}:PV^k(M)\to \PV^{k-1}(M)$ is identified with $\partial_{\xi_i}$. Now, we require our action function $S:M\to \R$ also be polynomial, and further, that $$S(x)=\frac{1}{2}\sum_{i,j} a_{ij}x^ix^j- b(x)$$ for $a_{ij}$ symmetric and non-degenerate and for $b\in I^3$ where $I=(x^1,...,x^n)\subset \K[[x^1,...,x^n]]$, that is, $b(x)$ a polynomial with no terms of degree less than 3. This implies that \begin{align*} \h\Delta_\mu & = \h\Delta_{\mu_0} -\iota_{dS} \\ & = \h \sum_i \partial_{x^i} \iota_{dx^i} - \sum_i (\partial_{x^i}S) \iota_{dx^i} \\ & = \h \sum_i \partial_{x^i} \iota_{dx^i} + \sum_i (\partial_{x^i}b) \iota_{dx^i} - \sum_{i,j}a_{ij}x^i\iota_{dx^j}\\ \end{align*} and under our identification this becomes $$\h\Delta_\mu = \h \sum_i \partial_{x^i}\partial_{\xi_i} + \sum_i (\partial_{x^i}b)\partial_{\xi_i} - \sum_{i,j}a_{ij}x^i\partial_{\xi_j}$$ The computation of the degree 0 cohomology of a given polynomial $f\in\K[[x_1,...,x_n]]$ under this differential is taken up in Gwilliam, Johnson-Freyd where it is shown the answer is precisely the Feynman diagram expansion for the expectation of $f$ which we expect.
Tuesday, March 18, 2014
Clifford Algebras and Spinors III: Bochner identity
Let $M$ be a Riemannian manifold, and let $Cl(M)$ be its Clifford bundle. Let $E \to M$ be any vector bundle with connection, and assume that $C^\infty(M, E)$ is a $Cl(M)$-module. We can define a Dirac operator $\mathcal{D}$ acting on sections of $E$ via the formula
\[ \mathcal{D} \sigma = \sum_{i=1}^n e_i \cdot \nabla_i \sigma \]
for any orthonormal frame $\{e_1, \dots, e_n\}$ on $M$, and where $\cdot$ denotes the Clifford module action. We demand that the connection on $E$ is compatible with Clifford multiplication in the following sense:
\[ \nabla_j (e_i \cdot \sigma) = (\nabla_j e_i) \cdot \sigma + e_i \cdot \nabla_j \sigma. \]
Let $R$ denote the curvature of $E$, i.e. we have
\[ [\nabla_i, \nabla_j] \sigma = R(e_i, e_j) \sigma+ \nabla_{[e_i, e_j]} \sigma \]
We can define an endomorphism $\mathcal{R}$ on $E$ by
\[ \mathcal{R} = \frac{1}{2} \sum_{ij} R(e_i, e_j). \]
Theorem. We have the identity $\mathcal{D}^2 = -\Delta + \mathcal{R}$.
Proof. We compute
\begin{align}
\mathcal{D}^2 \sigma &= \sum_{ij} e_i \nabla_i \left( e_j \nabla_j \sigma \right) \\
&= \sum_{ij} e_i e_j \nabla_i \nabla_j \sigma + e_i ( \nabla_i e_j ) \nabla_j \sigma \\
&= -\Delta \sigma + \frac{1}{2}\sum_{ij}e_i e_j [\nabla_i, \nabla_j] \sigma + \sum_{ij} e_i ( \nabla_i e_j) \nabla_j \sigma \\
&= -\Delta \sigma + \frac{1}{2}\sum_{ij}e_i e_j R(e_i, e_j) \sigma + \frac{1}{2}\sum_{ij} e_i e_j \nabla_{[e_i, e_j]} \sigma+ \sum_{ij} e_i ( \nabla_i e_j) \nabla_j \sigma \\
&= (-\Delta + \mathcal{R})\sigma + \frac{1}{2} \sum_{ij} \left( e_i e_j \nabla_{[e_i, e_j]}\sigma + e_i (\nabla_i e_j) \nabla_j + e_j (\nabla_j e_i) \nabla_i \right)\sigma
\end{align}
We will be done provided we can show that the last term vanishes. Notice that it is fully tensorial, since it can be expressed as $\mathcal{D}^2 + \Delta - \mathcal{R}$. On the other hand, the terms $[e_i, e_j]$ and $\nabla_j e_i$ are (by definition!) proportional to Christoffel symbols. Since we can always choose a frame so that these vanish at a point, these terms must vanish identically. Hence we have $0 = \mathcal{D}^2 + \Delta - \mathcal{R}$, as desired.
\[ \mathcal{D} \sigma = \sum_{i=1}^n e_i \cdot \nabla_i \sigma \]
for any orthonormal frame $\{e_1, \dots, e_n\}$ on $M$, and where $\cdot$ denotes the Clifford module action. We demand that the connection on $E$ is compatible with Clifford multiplication in the following sense:
\[ \nabla_j (e_i \cdot \sigma) = (\nabla_j e_i) \cdot \sigma + e_i \cdot \nabla_j \sigma. \]
Let $R$ denote the curvature of $E$, i.e. we have
\[ [\nabla_i, \nabla_j] \sigma = R(e_i, e_j) \sigma+ \nabla_{[e_i, e_j]} \sigma \]
We can define an endomorphism $\mathcal{R}$ on $E$ by
\[ \mathcal{R} = \frac{1}{2} \sum_{ij} R(e_i, e_j). \]
Theorem. We have the identity $\mathcal{D}^2 = -\Delta + \mathcal{R}$.
Proof. We compute
\begin{align}
\mathcal{D}^2 \sigma &= \sum_{ij} e_i \nabla_i \left( e_j \nabla_j \sigma \right) \\
&= \sum_{ij} e_i e_j \nabla_i \nabla_j \sigma + e_i ( \nabla_i e_j ) \nabla_j \sigma \\
&= -\Delta \sigma + \frac{1}{2}\sum_{ij}e_i e_j [\nabla_i, \nabla_j] \sigma + \sum_{ij} e_i ( \nabla_i e_j) \nabla_j \sigma \\
&= -\Delta \sigma + \frac{1}{2}\sum_{ij}e_i e_j R(e_i, e_j) \sigma + \frac{1}{2}\sum_{ij} e_i e_j \nabla_{[e_i, e_j]} \sigma+ \sum_{ij} e_i ( \nabla_i e_j) \nabla_j \sigma \\
&= (-\Delta + \mathcal{R})\sigma + \frac{1}{2} \sum_{ij} \left( e_i e_j \nabla_{[e_i, e_j]}\sigma + e_i (\nabla_i e_j) \nabla_j + e_j (\nabla_j e_i) \nabla_i \right)\sigma
\end{align}
We will be done provided we can show that the last term vanishes. Notice that it is fully tensorial, since it can be expressed as $\mathcal{D}^2 + \Delta - \mathcal{R}$. On the other hand, the terms $[e_i, e_j]$ and $\nabla_j e_i$ are (by definition!) proportional to Christoffel symbols. Since we can always choose a frame so that these vanish at a point, these terms must vanish identically. Hence we have $0 = \mathcal{D}^2 + \Delta - \mathcal{R}$, as desired.
Thursday, March 6, 2014
Clifford Algebras and Spinors, Part II: Spin Structures and Dirac Operators
A very good reference for today's material is Dan Freed's (unpublished) notes on Dirac operators, available here.
\[ (e_1 \cdots e_k)^t = e_k \cdots e_1, \ \beta(e_1 \dots e_k) = (-1)^k e_k \dots e_2 e_1 \]
There is a natural inclusion \(\mathbb E^n \hookrightarrow Cl(\mathbb E^n)\). Given \(x \in Cl(\mathbb E^n)\) and \(v \in \mathbb E^n\), we can consider the product \(x v x^t\). In general, this might not be contained in \(\mathbb E^n \subset Cl(\mathbb E^n)\).
Definition. We define the group \(Pin(n)\) to consist of all those \(g \in Cl(\mathbb E^n)\) such that
\[ g \beta(g) = 1, \ \ g v \beta(g) \subset \mathbb E^n \ \forall\ v \in \mathbb E^n. \]
Similarly, we define the group \(Spin(n)\) to be the subgroup of \(Pin(n)\) such that \(gg^t = 1\).
Theorem. The natural action of \(Pin(n)\) on \(\mathbb E^n\) is by othogonal transformations, giving a natural map \(Pin(n) \to O(n)\). This map is a double cover. Similarly, \(Spin(n)\) is a double cover of \(SO(n)\). If \(n \geq 2\), \(Spin(n)\) is simply connected.
The importance of the spin groups is due to the following basic fact. Suppose that \(G\) is a Lie group with Lie algebra \(\mathfrak{g}\). Any representation of \(G\) induces a representation of \(\mathfrak{g}\). However, given a representation of \(\mathfrak{g}\), it is not always possible to integrate it to a representation of \(G\). But it is always possible to integrate a representation of \(\mathfrak{g}\) to produce a representation of the universal cover of \(G\). For \(n \geq 2\), \(Spin(n)\) is the universal cover of \(SO(n)\).
Suppose that \(V\) is a representation of \(SO(n)\). Then we may form the associated bundle \(SO(M) \times_{SO(n)} V\), which is a vector bundle over \(M\) with structure group \(SO(n)\). If we take the defining representation then we obtain the tangent bundle, but of course there are many others. Unfortunately, since \(SO(n)\) is not simply connected, not every representation of \(\mathfrak{so}_n\) can be integrated to a representation of \(SO(n)\). At the level of geometry, this means that in a certain sense there are certain vector bundles over \(M\) that are "missing"! Even more disturbing, is that these "missing" bundles appear to be necessary to describe many of the fundamental particles that appear in the standard model--so this has real world consequences. The solution is to equip \(M\) with a spin structure.
Definition. A spin structure on \(M\) is a principal \(Spin(n)\)-bundle \(Spin(M)\) over \(M\) together with a bundle morphism \(Spin(M) \to SO(M)\) which is a reduction of structure (i.e., satisfies the obvious axioms).
As you might expect, not every manifold admits a spin structure, and spin structures may not be unique. Loosely speaking, a spin structure is a slightly stronger notion of orientability. Spin structures may always be chosen locally, and the obstruction to consistent gluing is not too difficult to characters as a certain \(\mathbb Z_2\) cohomology class, called the second Stiefel-Whitney class.
\[ S = Spin(M) \times_{Spin(n)} S_0 \]
which is called the spinor bundle. Moreover, since \(S_0\) is a Clifford module, there is well-defined notion of Clifford multiplication on sections of \(S\). We may then define the Dirac operator \(\mathcal{D}\) by
\[ \mathcal{D} = \sum_{a=1}^n c(e_a) \nabla_{e_a} \]
where \(\{e_a\}\) is any orthonormal frame, \(\nabla\) is the spin connection, and \(c\) denotes Clifford multiplication.
Next time: the Weitzenböck formula, and maybe a vanishing theorem.
Spin(n)
Consider the Clifford algebra \(Cl(\mathbb E^n)\) as constructed in yesterday's post. Define maps \(t, \beta: Cl(\mathbb E^n) \to Cl(\mathbb E^n)\) via\[ (e_1 \cdots e_k)^t = e_k \cdots e_1, \ \beta(e_1 \dots e_k) = (-1)^k e_k \dots e_2 e_1 \]
There is a natural inclusion \(\mathbb E^n \hookrightarrow Cl(\mathbb E^n)\). Given \(x \in Cl(\mathbb E^n)\) and \(v \in \mathbb E^n\), we can consider the product \(x v x^t\). In general, this might not be contained in \(\mathbb E^n \subset Cl(\mathbb E^n)\).
Definition. We define the group \(Pin(n)\) to consist of all those \(g \in Cl(\mathbb E^n)\) such that
\[ g \beta(g) = 1, \ \ g v \beta(g) \subset \mathbb E^n \ \forall\ v \in \mathbb E^n. \]
Similarly, we define the group \(Spin(n)\) to be the subgroup of \(Pin(n)\) such that \(gg^t = 1\).
Theorem. The natural action of \(Pin(n)\) on \(\mathbb E^n\) is by othogonal transformations, giving a natural map \(Pin(n) \to O(n)\). This map is a double cover. Similarly, \(Spin(n)\) is a double cover of \(SO(n)\). If \(n \geq 2\), \(Spin(n)\) is simply connected.
The importance of the spin groups is due to the following basic fact. Suppose that \(G\) is a Lie group with Lie algebra \(\mathfrak{g}\). Any representation of \(G\) induces a representation of \(\mathfrak{g}\). However, given a representation of \(\mathfrak{g}\), it is not always possible to integrate it to a representation of \(G\). But it is always possible to integrate a representation of \(\mathfrak{g}\) to produce a representation of the universal cover of \(G\). For \(n \geq 2\), \(Spin(n)\) is the universal cover of \(SO(n)\).
Spin Structures
Let \((M^n, g)\) be a Riemannian manifold. Recall that the frame bundle \(O(M)\) is the manifold consisting of pairs \((x, \mathbb{e})\) where \(x \in M\) and \(\mathbb{e} = \{e_1, \dots, e_n\}\) is an orthonormal frame in \(T_x M\). Since the orthogonal group \(O(n)\) acts on the set of orthonormal frames, this makes \(F(M)\) into a principal \(O(n)\) bundle over \(M\). Let us assume that \(M\) is oriented, so that we may reduce its structure group to \(SO(n)\).Suppose that \(V\) is a representation of \(SO(n)\). Then we may form the associated bundle \(SO(M) \times_{SO(n)} V\), which is a vector bundle over \(M\) with structure group \(SO(n)\). If we take the defining representation then we obtain the tangent bundle, but of course there are many others. Unfortunately, since \(SO(n)\) is not simply connected, not every representation of \(\mathfrak{so}_n\) can be integrated to a representation of \(SO(n)\). At the level of geometry, this means that in a certain sense there are certain vector bundles over \(M\) that are "missing"! Even more disturbing, is that these "missing" bundles appear to be necessary to describe many of the fundamental particles that appear in the standard model--so this has real world consequences. The solution is to equip \(M\) with a spin structure.
Definition. A spin structure on \(M\) is a principal \(Spin(n)\)-bundle \(Spin(M)\) over \(M\) together with a bundle morphism \(Spin(M) \to SO(M)\) which is a reduction of structure (i.e., satisfies the obvious axioms).
As you might expect, not every manifold admits a spin structure, and spin structures may not be unique. Loosely speaking, a spin structure is a slightly stronger notion of orientability. Spin structures may always be chosen locally, and the obstruction to consistent gluing is not too difficult to characters as a certain \(\mathbb Z_2\) cohomology class, called the second Stiefel-Whitney class.
Spin Connection and Dirac Operators
The reduction of structure \(Spin(M) \to SO(M)\) allows us to pull back the Levi-Civita connection on \(SO(M)\) to obtain a connection on \(Spin(M)\), called the spin connection. Let \(S_0\) be the spinor module described in the previous post. Then we may construct the associated bundle\[ S = Spin(M) \times_{Spin(n)} S_0 \]
which is called the spinor bundle. Moreover, since \(S_0\) is a Clifford module, there is well-defined notion of Clifford multiplication on sections of \(S\). We may then define the Dirac operator \(\mathcal{D}\) by
\[ \mathcal{D} = \sum_{a=1}^n c(e_a) \nabla_{e_a} \]
where \(\{e_a\}\) is any orthonormal frame, \(\nabla\) is the spin connection, and \(c\) denotes Clifford multiplication.
Next time: the Weitzenböck formula, and maybe a vanishing theorem.
Wednesday, March 5, 2014
Clifford Algebras and Spinors
Clifford Algebras
Today I'd like to write some brief notes about Clifford algebras and spinors. A classic reference is the paper "Clifford Modules" by Atiyah-Bott-Shapiro. Clifford algebras not only useful in algebra and geometry, but are essential for the construction of theories with fermions. Let \(V\) be a vector space with a non-degenerate symmetric bilinear form \(B\). We define the Clifford algebra \(Cl(V, B)\) to be the unital associative algebra generated by \(v \in V\) subject to the relation\[ vw + wv = -2B(v,w) \]
Equivalently, the definiting relation is \(v^2 = -B(v,v)\).
The Clifford algebra inherits a \(\mathbb Z\)-filtration as well as a \(\mathbb Z_2\)-grading from the tensor algebra. In fact, we have an analogue of the Poincare-Birkhoff-Witt theorem for Lie algebras:
Theorem The associated graded algebra of \(Cl(V,B)\) is naturally isomorphic to the exterior algebra on \(V\).
In this way, we may view the Clifford algebra as a quantization of the exterior algebra, much in the same way that \(U(\mathfrak g)\) is a quantization of the Poisson algebra of functions on \(\mathfrak g^\ast\) for a Lie algebra \(\mathfrak g\).
Example. Take (V,B) to be the Euclidean space \(\mathbb E^1\). Then we have a single generator \(e\) satisfying the relation \(e^2 = -1\). Hence
\[ Cl(\mathbb R) \cong \mathbb R \cdot 1 \oplus \mathbb R \cdot e \cong \mathbb C \]
Where the isomorpism is given by \(e \mapsto i = \sqrt{-1}\).
Example. Take \(\mathbb E^2\). We have generators \(e_1, e_2\) both squaring to -1, and additionally we have \(e_1 e_2 = e_2 e_1\). We can define an isomorphism from \(Cl(\mathbb E^2)\) to the quaternions \(\mathbb H\) by \(e_1 \mapsto i, e_2 \mapsto j\).
Spinors
Now consider the complexified Clifford algebra, denoted \(\mathbb{C}l(V)\). Since we can now take square roots of negative numbers, the complex Clifford algebra is insensitive to the signature (as long as our bilinear form is non-degenerate). Denote by \(C_n\) the complex Clifford algebra \(Cl(\mathbb C^n)\),where \(\mathbb C^n\) is equipped with the standard bilinear form \((x,y) = \sum_{i=1}^n x_i y_i\).
Definition. A subspace \(W \subset \mathbb C^n\) is isotropic if the restriction of the standard bilinear form to \(W\) is identically 0. A maximal isotropic subspace is an isotropic subspace that is not properly contained in any other isotropic subspace.
Theorem. Let \(W\) be a maximal isotropic subspace, and let\( \{w_1, \dots, w_k\}\) be a basis of \(W\). Let \(\omega = w_1 \cdots w_k \in C_n\), and let \(S = C_n \cdot \omega\). If n is even, then \(S\) is an irreducible Clifford module. If n is odd, then \(S=S^+ \oplus S^-\) is a direct sum irreducible Clifford modules, and \(S^+ \cong S^-\).
Irreducible Clifford modules are called spinor modules. This description of spinor modules allows one to prove straightforwardly the following complete classification of complex Clifford algebras.
Corollary. We have \(C_{2m} \cong \mathrm{End}(\mathbb C^m)\) and \(C_{2m+1} \cong \mathrm{End}(\mathbb C^m) \oplus \mathrm{End}(\mathbb C^m)\).
Note that this classification depends on n mod 2, which is closely related to Bott periodicity. There is a similar classification of real Clifford algebras.
Dirac Operators
Now we come to the real importance of Clifford algebras. Consider Euclidean space \(\mathbb{E}^n\) and let \(S\) be a spinor module for its Clifford algebra. We define the Dirac operator acting on \(S\)-valued functions as\[ D f = \sum_{i=1}^n e_i \cdot \partial_i f \]
Now the amazing property of \(D\) is the following:
\[ D^2 = \sum_{i,j} e_i e_j \partial_i \partial_j = \sum_i e_i^2 \partial_i^2 + \sum_{i,j} e_i e_j [\partial_i, \partial_j] = -\Delta \]
hence the Dirac operator provides an algebraic (as opposed to pseudodifferential) square root of the Laplacian.
To Be Added in an Update...
Supersymmetric point particle, Dirac operators on spin manifolds, Weitzenböck formula, spinor reps of Lorentz algebra, N=1 susy.Sunday, February 23, 2014
Virasoro Algebra
Conformal Invariance in 2D
To begin, recall that in two dimensions, the conformal transformations are generated by holomorphic and anti-holomorphic transformations. At the infinitesimal level, let \(\ell_n := -z^{n+1} \partial_z\) be a basis of holomorphic vector fields. These satisfy the Witt algebra\[ [\ell_m, \ell_n] = (m-n)\ell_{m+n}. \]
Similarly, we can define \(\bar{\ell}_m = -\bar{z}^{n+1} \partial_{\bar{z}}\), and in addition to the Witt algebra these new generators satisfy \([\bar{\ell}_m, \ell_n]=0\).
Now, we could try to define a 2D conformal quantum field theory to be a unitary representation of the Witt algebra (or rather, of two copies of the Witt algebra, since we have both holomorphic and anti-holomorphic vector fields--but nevermind that). But this is too naive.
Central Extensions
Recall that in quantum mechanics, states are represented by vectors in some Hilbert space \(\mathcal{H}\). However, the state \(|\phi\rangle\) and \(\alpha|\phi\rangle\) are physically equivalent for any non-zero complex number \(\alpha\). The reason, of course, is that the expectation value of an operator \(\mathcal{O}\) is defined to be \(\langle \phi|\mathcal{O}|\phi\rangle / \langle \phi|\phi\rangle\), and such expressions are invariant under rescaling in \(\mathcal{H}\).Thus, a symmetry group \(G\) for a theory does not necessarily act via a map \(G \to U(\mathcal{H})\). It suffices to have a projective representation \(G \to PU(\mathcal{H})\). Let \(\mathfrak{g}, \mathfrak{pu}\) be the Lie algebras of \(G\) and \(PU\), respectively. A projective representation gives a map
\[ \mathfrak{g} \to \mathfrak{pu}. \]
Since \(PU\) is a quotient of \(U\), we have a short exact sequence
\[ 0 \to \mathbb{C} \to \mathfrak{u} \to \mathfrak{pu} \to 0. \]
Now let \(\hat{\mathfrak{g}}\) be defined as
\[ \hat{\mathfrak{g}} = \{ (\xi, \eta) \in \mathfrak{u}\oplus\mathfrak{g} \ | \ \pi(\xi) = \rho(\eta) \} \]
This comes with a natural projection \(\hat{\mathfrak{g}} \to \mathfrak{g}\). If we suppose that the projective representation \(\rho\) is faithful, then the kernel of this map is exactly \(\mathbb{C}\). Hence, a faithful projective representation of \(\mathfrak{g}\) yields a short exact sequence of Lie algebras
\[ 0 \to \mathbb{C} \to \hat{\mathfrak{g}} \to \mathfrak{g} \to 0. \]
We have obtained a central extension of \(\mathfrak{g}\).
Virasoro Algebra
Finally, we can define the Virasoro algebra. It has generators \(L_n\) and \(c\), with defining relations\[ [L_m, L_n] = (m-n) L_{m+n} + \frac{c}{12}(m^3-m) \delta_{m+n,0}, [c, L_n] = 0. \]
The generator \(c\) acts as a scalar in any irreducible representation, and its value is called the central charge. The factor of \(1/12\) is entirely conventional. Now, the amazing fact is the following.
Theorem. Up to equivalence, the Virasoro algebra is the unique non-trivial central extension of the Witt algebra.
Proof sketch. This is essentially just a calculation. Any central extension has to be of the form
\[ [L_m, L_n] = (m-n) L_{m+n} + A(m,n) c \]
for some function \(A(m,n)\). If we make the replacement \(L_m \mapsto L_m + a_m c\), then we have
\[ [L_m, L_n] = (m-n) L_{m+n} + \left( A(m,n) + (m-n) a_{m+n} \right) c \]
Taking \(n = 0\), we have
\[ [L_m, L_0] = m L_{m} + \left( A(m,0) + m a_{m} \right) c \]
Hence for \(m\neq0\) we can take \(a_m = m^{-1} A(m,0)\). Having done this, we are now free to assume that \(A(m,0) = 0 \) for all \(m\). Then we may apply the Jacobi identity to deduce that \(A(m,n)=0\) except possibly for \(m=-n\), so that \(A(m,n)\) can be written in the form \(A(m,n) = A_m \delta_{m+n, 0}\). Finally, another application of the Jacobi identity yields a simple recurrence relation for the coefficients \(A_m\), and it is easily seen that every solution of this recurrence is proportional to \(m^3-m\).
Now we can take our (preliminary, and still too naive) definition of a quantum conformal field theory to be a unitary representation of the Virasoro algebra.
Stress-Energy Tensor and OPE
The operator \(L_0\) behaves like the Hamiltonian of the theory, and the Virasoro relations show that \(L_n\) for \(n>0\) act as lowering operators. Hence, in a physically sensible representation, the vacuum vector \(|\Omega\rangle\) will be annihilated by \(L_n\) for all \(n > 0\). Unitary requires \(L_n^\dagger = L_{-n}\), so additionally we have \(\langle \Omega|L_n = 0\) for \(n < 0\). Hence
\[ \langle \Omega | L_m L_n | \Omega \rangle = 0 \ \textrm{unless}\ n \leq 0, m \geq 0 \]
Now define the stress-energy tensor to be the operator-valued formal power series
\[ T(z) = \sum_n \frac{L_n}{z^{n+2}} \]We can consider the vacuum expectation of the product \(T(z) T(w)\). By the above remarks, many terms in the expansion will vanish. In fact, it is a straightforward (but tedious!) exercise to check the following.
Theorem. The stress-energy tensor satisfies the operator product expansion
\[ T(z) T(w) \sim \frac{c/2}{(z-w)^4} + \frac{2 T(w)}{(z-w)^2} + \frac{\partial_w T(w)}{z-w} \]
where \(\sim\) denotes that the left- and right-hand sides are equal up to the addition of terms with vanishing vev and/or regular as \(z \to w\).
Friday, December 28, 2012
BRST and Lie Algebra Cohomology
We saw in previous posts that gauge-fixing is intimately related to BRST cohomology. Today I want to explain the underlying mathematical formalism, as it is actually something very well-known: Lie algebra cohomology. Let \(\mathfrak{g}\) be a Lie algebra and \(M\) a \(\mathfrak{g}\)-module. We will construct a cochain complex that computes the Lie algebra cohomology with values in \(M\), \(H^i(\mathfrak{g}, M)\). Out of thin air, we define
\[ C^\ast(\mathfrak{g}, M) = M \otimes \wedge^\ast \mathfrak{g}^\ast. \]
The grading is just the grading induced by the grading on \(\wedge^\ast \mathfrak{g}^\ast\), which we identify with the BRST ghost number. Let \(e_i\) be a basis for \(M\) and \(T_a\) be a basis for \(\mathfrak{g}\), with canonical dual basis \(S^a\). The differential is defined on generators to be
\[ d e_i = \rho(T_a) e_i \otimes S^a \]
\[ d S^a = \frac{1}{2} f^a_{bc} S^b \wedge S^c \]
where \(\rho: \mathfrak{g} \to \mathrm{End}(M)\) is the representation and \(f^a_{bc}\) are the structure constants of the group. This differential is then extended to satisfy the graded Leibniz rule, and is easily verified to satisfy \(d^2 = 0\) (this is just the Jacobi identity). The Lie algebra cohomology is just the cohomology of this cochain complex. Essentially by definition, we see that
\[ H^0(\mathfrak{g}, M) = \{m \in M \ | \ \xi \cdot m = 0 \ \forall \ \xi \in \mathfrak{g} \}, \]
i.e. \(H^0(\cdot) = (\cdot)^\mathfrak{g}\) is the invariants functor. In fact, this can be taken to be the defining property of Lie algebra cohomology:
Theorem \(H^k(\mathfrak{g}, M) = R^k (M)^\mathfrak{g}\).
Returning to field theory, we see (modulo some hard technicalities!) that, roughly, \(\mathfrak{g}\) is the Lie algebra of infinitesimal gauge transformations, and \(M\) is the algebra of functions on the space of all connections. The ghost and anti-ghost fields can then be seen to be the multiplication and contraction operators. To wit, we can take \(c^a\) to be the operator
\[ c^a: f \mapsto S^a \wedge f \]
and take \(\bar{c}^a\) to be the operator
\[ \bar{c}^a: f \mapsto \frac{\partial}{\partial S^a} f = T_a \lrcorner f.\]
Then we have
\[ [c^a, \bar{c}^b] = \delta^{ab} \]
so that \(\bar{c}\) is indeed the antifield of \(c\).
\[ C^\ast(\mathfrak{g}, M) = M \otimes \wedge^\ast \mathfrak{g}^\ast. \]
The grading is just the grading induced by the grading on \(\wedge^\ast \mathfrak{g}^\ast\), which we identify with the BRST ghost number. Let \(e_i\) be a basis for \(M\) and \(T_a\) be a basis for \(\mathfrak{g}\), with canonical dual basis \(S^a\). The differential is defined on generators to be
\[ d e_i = \rho(T_a) e_i \otimes S^a \]
\[ d S^a = \frac{1}{2} f^a_{bc} S^b \wedge S^c \]
where \(\rho: \mathfrak{g} \to \mathrm{End}(M)\) is the representation and \(f^a_{bc}\) are the structure constants of the group. This differential is then extended to satisfy the graded Leibniz rule, and is easily verified to satisfy \(d^2 = 0\) (this is just the Jacobi identity). The Lie algebra cohomology is just the cohomology of this cochain complex. Essentially by definition, we see that
\[ H^0(\mathfrak{g}, M) = \{m \in M \ | \ \xi \cdot m = 0 \ \forall \ \xi \in \mathfrak{g} \}, \]
i.e. \(H^0(\cdot) = (\cdot)^\mathfrak{g}\) is the invariants functor. In fact, this can be taken to be the defining property of Lie algebra cohomology:
Theorem \(H^k(\mathfrak{g}, M) = R^k (M)^\mathfrak{g}\).
Returning to field theory, we see (modulo some hard technicalities!) that, roughly, \(\mathfrak{g}\) is the Lie algebra of infinitesimal gauge transformations, and \(M\) is the algebra of functions on the space of all connections. The ghost and anti-ghost fields can then be seen to be the multiplication and contraction operators. To wit, we can take \(c^a\) to be the operator
\[ c^a: f \mapsto S^a \wedge f \]
and take \(\bar{c}^a\) to be the operator
\[ \bar{c}^a: f \mapsto \frac{\partial}{\partial S^a} f = T_a \lrcorner f.\]
Then we have
\[ [c^a, \bar{c}^b] = \delta^{ab} \]
so that \(\bar{c}\) is indeed the antifield of \(c\).
Labels:
brst-bv,
gauge theory,
path integral,
quantum field theory
Sunday, December 23, 2012
BRST
Finally, I want to discuss gauge-invariant of the gauge-fixed theory. (!?) We saw in the previous posts that if we have a gauge theory with connection \(A\) and matter fields \(\psi\), in order to derive sensible Feynman rules we have to introduce a gauge-fixing function \(G\) as well as Fermionic fields \(c, \bar{c}\), the ghosts. (Note: last time I used \(\eta, \bar{\eta}\) for the ghosts but I want to match the more standard notation, so I've switched to \(c, \bar{c}\)).
Usually it is convenient to use the gauge-fixing function \(G(A) = \partial^\mu A_\mu\). Under an infinitesimal gauge-transformation \(\lambda\), \(A\) transforms as
\[ A \mapsto -\nabla \lambda, \]
so \(G(A)\) transforms as
\[ G(A) \mapsto G(A) - \partial^\mu \nabla_\mu \lambda. \]
Hence the term in the Lagrangian involving the ghosts is
\[ -\bar{c}^a \partial^\mu \nabla_\mu^{ab} c^b, \]
and our gauge-fixed Lagrangian is
\[ \mathcal{L} = -\frac{1}{4} |F|^2 + \bar{\psi}(iD\!\!\!/-m)\psi + -\frac{|\partial^\mu A_\mu|^2}{2\xi}
- \bar{c}^a \nabla_\mu^{ab} c^b \]
Introducing an auxiliary filed \(B^a\), this is of course equivalent to
\[ \mathcal{L} = -\frac{1}{4} |F|^2 + \bar{\psi}(iD\!\!\!/-m)\psi + \frac{\xi}{2} B^a B_a
+ B^a \partial^\mu A_{\mu a} - \bar{c}^a \nabla_\mu^{ab} c^b. \]
Now, there are two questions one might ask: (1) how can we tell that this is a gauge-theory? i.e., what remains of the original gauge symmetry? and (2) does the resulting theory depend in any way on the choice of gauge-fixing function?
The answer to both of these questions is BRST symmetry. The field \(c\) is Lie-algebra valued, so we could think of it as being an infinitesimal gauge transformation. Rather, for \(\epsilon\) a constant odd variable, \(\epsilon c\) is even and an honest infinitesimal gauge transformation. Under this transformation, we have
\[ \delta_\epsilon A = -\nabla (\epsilon c) = -\epsilon \nabla c. \]
Then we define a graded derivation \(\delta\) by
\[ \delta A = - \nabla c. \]
We have a grading by ghost number, where \(\mathrm{gh}(A) = 0, \mathrm{gh}(\psi) = 0, \mathrm{gh}(c) = 1, \mathrm{gh}(\bar{c}) = -1\). We would like to extend \(\delta\) to a derivation of degree \(+1\) that squares to 0. First, we should figure out what \(\delta c\) is. We compute:
\begin{align}
0 &= \delta^2 A \\
&= \delta(-\nabla c) \\
&= -\partial \delta c - (\delta A) c - A (\delta c) + (\delta c) A - c (\delta A) \\
&= -\partial \delta c + (\nabla c) c + c (\nabla c) - [A, \delta c] \\
&= -\nabla(\delta c) + \nabla(c^2).
\end{align}
From this, we see that \(\nabla(\delta c) = \nabla(c^2)\), so we can set
\[ \delta c = c^2 = \frac{1}{2}[c, c]. \]
Then \(\delta^2 c = 0\) is just the Jacobi identity for the group's Lie algebra! Finally, we would like to extend \(\delta\) to act on \(\psi\), \(B\), and \(\bar{c}\) so that \(\delta \mathcal{L} = 0\), and \(\delta^2 = 0\). Since the action on \(A\) is by infinitesimal gauge transformation, this leaves the curvature term of \(\mathcal{L}\) invariant. Similarly, the \(\psi\) term is invariant if we simply take
\[ \delta \psi = c \cdot \psi \]
where dot denotes the infinitesimal gauge transformation. Using the known rules for \(\delta\), we find that
\[ \delta \mathcal{L} = \frac{\xi}{2} \left(\delta B B + B \delta B \right) + \delta B \cdot \partial^\mu A_\mu
- B \cdot \partial^\mu \nabla_\mu c - \delta\bar{c} \cdot \partial^\mu \nabla_\mu c \]
By comparing coefficients, we find (together with what we've already computed)
\begin{align}
\delta A &= -\nabla c \\
\delta \psi &= c \cdot \psi \\
\delta c &= \frac{1}{2}[c,c] \\
\delta \bar{c} &= B \\
\delta B &= 0.
\end{align}
This is the BRST differential. Now, suppose that \(\mathcal{O}(A, \psi)\) is a local operator involving the physical fields \(A\) and \(psi\). Then by construction,\(delta O\) is the change of \(O\) under an infinitesimal gauge transformation. Hence, we find
\[ \langle \delta \mathcal{O} \rangle = 0 \]
for any local observable \(\mathcal{O}\). This just follows from integration by parts (this is where we have to assume the measure is \(\delta\)-closed). Now, why is this significant? First, this tells us that the space of physical observables is
\[ H^0(C^\ast_{\mathrm{BRST}}, \delta) \]
where \(C^\ast_{\mathrm{BRST}}\) is the cochain complex of local observables, graded by ghost number.
Now, the real power of the BRST formalism is the following. We find that the gauge-fixed Lagrangian can be written as
\[ \mathcal{L}_{gf} = \mathcal{L}_0 +\delta \left(\bar{c} \frac{B}{2} + \bar{c}\Lambda\right) \]
where \( \Lambda = \partial^\mu \nabla_\mu A \) is our gauge-fixing function, and \(\mathcal{L}_0\) is the original Lagrangian without gauge-fixing. Now the point is, any two choices of gauge fixing differ by terms which are BRST exact, and hence give the same expectation values on the physical observables \(H^0\). So we have restored gauge invariance, while obtaining a gauge-fixed perturbation theory!
Usually it is convenient to use the gauge-fixing function \(G(A) = \partial^\mu A_\mu\). Under an infinitesimal gauge-transformation \(\lambda\), \(A\) transforms as
\[ A \mapsto -\nabla \lambda, \]
so \(G(A)\) transforms as
\[ G(A) \mapsto G(A) - \partial^\mu \nabla_\mu \lambda. \]
Hence the term in the Lagrangian involving the ghosts is
\[ -\bar{c}^a \partial^\mu \nabla_\mu^{ab} c^b, \]
and our gauge-fixed Lagrangian is
\[ \mathcal{L} = -\frac{1}{4} |F|^2 + \bar{\psi}(iD\!\!\!/-m)\psi + -\frac{|\partial^\mu A_\mu|^2}{2\xi}
- \bar{c}^a \nabla_\mu^{ab} c^b \]
Introducing an auxiliary filed \(B^a\), this is of course equivalent to
\[ \mathcal{L} = -\frac{1}{4} |F|^2 + \bar{\psi}(iD\!\!\!/-m)\psi + \frac{\xi}{2} B^a B_a
+ B^a \partial^\mu A_{\mu a} - \bar{c}^a \nabla_\mu^{ab} c^b. \]
Now, there are two questions one might ask: (1) how can we tell that this is a gauge-theory? i.e., what remains of the original gauge symmetry? and (2) does the resulting theory depend in any way on the choice of gauge-fixing function?
The answer to both of these questions is BRST symmetry. The field \(c\) is Lie-algebra valued, so we could think of it as being an infinitesimal gauge transformation. Rather, for \(\epsilon\) a constant odd variable, \(\epsilon c\) is even and an honest infinitesimal gauge transformation. Under this transformation, we have
\[ \delta_\epsilon A = -\nabla (\epsilon c) = -\epsilon \nabla c. \]
Then we define a graded derivation \(\delta\) by
\[ \delta A = - \nabla c. \]
We have a grading by ghost number, where \(\mathrm{gh}(A) = 0, \mathrm{gh}(\psi) = 0, \mathrm{gh}(c) = 1, \mathrm{gh}(\bar{c}) = -1\). We would like to extend \(\delta\) to a derivation of degree \(+1\) that squares to 0. First, we should figure out what \(\delta c\) is. We compute:
\begin{align}
0 &= \delta^2 A \\
&= \delta(-\nabla c) \\
&= -\partial \delta c - (\delta A) c - A (\delta c) + (\delta c) A - c (\delta A) \\
&= -\partial \delta c + (\nabla c) c + c (\nabla c) - [A, \delta c] \\
&= -\nabla(\delta c) + \nabla(c^2).
\end{align}
From this, we see that \(\nabla(\delta c) = \nabla(c^2)\), so we can set
\[ \delta c = c^2 = \frac{1}{2}[c, c]. \]
Then \(\delta^2 c = 0\) is just the Jacobi identity for the group's Lie algebra! Finally, we would like to extend \(\delta\) to act on \(\psi\), \(B\), and \(\bar{c}\) so that \(\delta \mathcal{L} = 0\), and \(\delta^2 = 0\). Since the action on \(A\) is by infinitesimal gauge transformation, this leaves the curvature term of \(\mathcal{L}\) invariant. Similarly, the \(\psi\) term is invariant if we simply take
\[ \delta \psi = c \cdot \psi \]
where dot denotes the infinitesimal gauge transformation. Using the known rules for \(\delta\), we find that
\[ \delta \mathcal{L} = \frac{\xi}{2} \left(\delta B B + B \delta B \right) + \delta B \cdot \partial^\mu A_\mu
- B \cdot \partial^\mu \nabla_\mu c - \delta\bar{c} \cdot \partial^\mu \nabla_\mu c \]
By comparing coefficients, we find (together with what we've already computed)
\begin{align}
\delta A &= -\nabla c \\
\delta \psi &= c \cdot \psi \\
\delta c &= \frac{1}{2}[c,c] \\
\delta \bar{c} &= B \\
\delta B &= 0.
\end{align}
This is the BRST differential. Now, suppose that \(\mathcal{O}(A, \psi)\) is a local operator involving the physical fields \(A\) and \(psi\). Then by construction,\(delta O\) is the change of \(O\) under an infinitesimal gauge transformation. Hence, we find
An operator \(\mathcal{O}\) is gauge invariant \(\iff \delta\mathcal{O} = 0\).Now, suppose the functional measure \(\mathcal{D}A \mathcal{D}\psi \mathcal{D}B \mathcal{D}c \mathcal{D}\bar{c}\) is gauge-invariant, i.e. is BRST closed. (This assumption is equivalent to the absence of anomalies, but we'll completely ignore this in today's post.) Then we have
\[ \langle \delta \mathcal{O} \rangle = 0 \]
for any local observable \(\mathcal{O}\). This just follows from integration by parts (this is where we have to assume the measure is \(\delta\)-closed). Now, why is this significant? First, this tells us that the space of physical observables is
\[ H^0(C^\ast_{\mathrm{BRST}}, \delta) \]
where \(C^\ast_{\mathrm{BRST}}\) is the cochain complex of local observables, graded by ghost number.
Now, the real power of the BRST formalism is the following. We find that the gauge-fixed Lagrangian can be written as
\[ \mathcal{L}_{gf} = \mathcal{L}_0 +\delta \left(\bar{c} \frac{B}{2} + \bar{c}\Lambda\right) \]
where \( \Lambda = \partial^\mu \nabla_\mu A \) is our gauge-fixing function, and \(\mathcal{L}_0\) is the original Lagrangian without gauge-fixing. Now the point is, any two choices of gauge fixing differ by terms which are BRST exact, and hence give the same expectation values on the physical observables \(H^0\). So we have restored gauge invariance, while obtaining a gauge-fixed perturbation theory!
Fadeev-Popov Ghosts, continued
Last time I sketched how we can represent an integral over a submanifold \(M \subset \mathbb{R}^n\) by an integral of the form
\[ \int_{\mathbb{R}^n} f(x) \delta(G(x)) \exp\left(\bar{\eta}G(x+\eta) \right) d\eta d\bar{\eta} dx. \]
Here, \(\eta, \bar{\eta}\) are Fermionic variables called Fadeev-Popov ghosts, which are introduced to cancel an unwanted determinant factor. The function \(G(x)\) singles out the submanifold \(M\) as \(M = G^{-1}(0)\).
Now suppose that we start with a vector (or affine) space \(V\), which is acted on by a group \(H\). We would like to undstand integrals over the quotient \(V / H\) in terms of integrals over \(V\). Suppose there is some function \(G(x)\) on \(V\) satisfying the following property:
We have some integral
\[Z = \int \delta(G(x) - w) \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta\right\} dx d\eta d\bar{\eta} \]
which is independent of \(w\). So we add a Gaussian weight an integrate over \(w\):
\begin{align}
Z' &= \int \delta(G(x) -w) \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta - \frac{1}{2\xi} |w|^2\right\}
dx dw d\eta d\bar{\eta}\\
&= \int \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta - \frac{1}{2\xi} |G(x)|^2 \right\} dx d\eta d\bar{\eta}.
\end{align}
Here, \(\xi\) is an arbitrary real positive constant, and we denote the new integral by \(Z'\) to indicate that it differs from the old path integral \(Z\) by (at most) an overall constant. Now the important thing is that the new action appear in the integrand of \(Z'\) is gauge-fixed and hence there is no problem whatsoever in deriving sensible, meaningful Feynman rules. The gauge-fixing term \(|G(x)|^2\) serves to make the action non-degenerate, so that propagators are well-defined, while the term involving the Fermions \(\eta, \bar{\eta}\) generates new Feynman rules that "cancel" the superfluous degrees of freedom due to gauge redundancy.
The question remains, what if we choose some other gauge-fixing function? i.e., what happens if we perturb \(G(x)\) to some new function satisfying property (*)? We'll answer this using the BRST formalism.
\[ \int_{\mathbb{R}^n} f(x) \delta(G(x)) \exp\left(\bar{\eta}G(x+\eta) \right) d\eta d\bar{\eta} dx. \]
Here, \(\eta, \bar{\eta}\) are Fermionic variables called Fadeev-Popov ghosts, which are introduced to cancel an unwanted determinant factor. The function \(G(x)\) singles out the submanifold \(M\) as \(M = G^{-1}(0)\).
Now suppose that we start with a vector (or affine) space \(V\), which is acted on by a group \(H\). We would like to undstand integrals over the quotient \(V / H\) in terms of integrals over \(V\). Suppose there is some function \(G(x)\) on \(V\) satisfying the following property:
For each level \(w\) of \(G\), the subspace \(M_w := G^{-1}(w)\) intersects the orbits transversely, and furthermore every \(H\)-orbit intersects \(M_w\) exactly once. (*)We call such a function a gauge-fixing function, and a level \(w\) a gauge-fixing. By assumption, we have \(V/H \cong M_w\) for any \(w\). Hence using the integral we derived last time, we can integrate over \(M_w\) for any particular choice of \(w\), and this ought to be the same as integrating over \(V/H\). The problem, however, is that in the QFT setting it's not clear what the Feynman rules should be for such a path integral. The final trick is that since the answer should be independent of \(w\), and by integrating over all possible \(w\) we obtain a Lagrangian from which we can derive sensible Feynman rules.
We have some integral
\[Z = \int \delta(G(x) - w) \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta\right\} dx d\eta d\bar{\eta} \]
which is independent of \(w\). So we add a Gaussian weight an integrate over \(w\):
\begin{align}
Z' &= \int \delta(G(x) -w) \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta - \frac{1}{2\xi} |w|^2\right\}
dx dw d\eta d\bar{\eta}\\
&= \int \exp\left\{\frac{i}{\hbar}S(x) + \bar{\eta}dG \eta - \frac{1}{2\xi} |G(x)|^2 \right\} dx d\eta d\bar{\eta}.
\end{align}
Here, \(\xi\) is an arbitrary real positive constant, and we denote the new integral by \(Z'\) to indicate that it differs from the old path integral \(Z\) by (at most) an overall constant. Now the important thing is that the new action appear in the integrand of \(Z'\) is gauge-fixed and hence there is no problem whatsoever in deriving sensible, meaningful Feynman rules. The gauge-fixing term \(|G(x)|^2\) serves to make the action non-degenerate, so that propagators are well-defined, while the term involving the Fermions \(\eta, \bar{\eta}\) generates new Feynman rules that "cancel" the superfluous degrees of freedom due to gauge redundancy.
The question remains, what if we choose some other gauge-fixing function? i.e., what happens if we perturb \(G(x)\) to some new function satisfying property (*)? We'll answer this using the BRST formalism.
Friday, December 21, 2012
Fadeev-Popov Ghosts
Today I want to review the Fadeev-Popov procedure, with a view toward BRST and eventually BV.
First we'll review the Fedeev-Popov method, to motivate the introduction of ghosts. Suppose we have a gauge theory involving a \(G\)-connection \(A\) and some field \(\phi\) charged under \(G\). Under a gauge transformation \(g(x)\), \(\phi\) transforms as
\[ \phi \mapsto g \cdot \phi. \]
We would like that the covariant derivative transforms in the same way, i.e.
\[ \nabla \phi \mapsto g \nabla \phi. \]
In terms of the connection 1-form \(A\), the covariant derivative is
\[ \nabla = d + A. \]
Let \(\nabla'\) denote the gauge-transformed covariant derivative, and \(\phi'\) the gauge-transformed field. Then we want
\[ \nabla' \phi' = g \nabla \phi. \]
We compute
\begin{align}
\nabla' \phi' &= (d + A')(g \phi) \\
&= dg \phi + g d\phi + A'(g\phi) \\
&= gg^{-1}dg \phi + gd\phi + g g^{-1} A' g\phi \\
&= g(d\phi + g^{-1} A' g \phi + g^{-1} dg \phi \\
&= g(d\phi + A\phi)
\end{align}
Comparing terms, we see that
\[ A = g^{-1} A' g + g^{-1} dg, \]
so upon re-arranging we have
\[ A' = g A g^{-1} - dg g^{-1} \]
This causes a problem: at critical points of the action, the Hessian of the action is degenerate in directions tangent to the gauge orbits. This means that the propagator is not well-defined, and there is no obvious way to derive the Feynman rules for perturbation theory. The solution is to take the quotient by gauge-transformations. To do this, we pick some gauge-fixing function \(G(A)\) which ought to be transverse to the orbits. Then we can restrict to the space \(G(A) = 0\), on which the Hessian of the action is non-degenerate, leading to a well-defined propagator. Formally, the path integral is
\[ Z = \int_{\{G(A) = 0\}} \exp{\frac{i}{\hbar} S[A, \phi]} \mathcal{D}A \mathcal{D}\phi \]
Formally, this suggests that the path integral should be something like
\[ Z = \int \delta(G(A)) \exp{\frac{i}{\hbar} S[A, \phi]} \mathcal{D}A \mathcal{D}\phi, \]
but this is not quite right! To understand the source of the problem, we'll first study the finite-dimensional case and then use this to solve the problem in infinite-dimensions.
Suppose we are on \(\mathbb{R}^n\), and we would like to integrate a function \(f(x)\) over a submanifold \(M\) defined by \(M = G^{-1}(0)\) for some smooth function \(G: \mathbb{R}^n \to \mathbb{R}^k\). Naively, we might expect that the answer is
\[ \int_M f(x) \stackrel{?}{=} \int f(x) \delta(G(x)) dx. \]
To see why this is not correct, write the delta function as
\[ \delta(G(x)) = \frac{1}{(2\pi)^k}\int e^{ip\cdot G(x)} d^k p. \]
We can regularize this by taking the limit as \(\epsilon \to 0\) of
\[ \frac{1}{(2\pi)^k}\int \exp\left\{ip\cdot G(x) -\frac{\epsilon}{2} |p|^2\right\} d^k p \]
This integral is Gaussian, so we obtain explicitly
\[ \left(\frac{2\pi}{\epsilon} \right)^{\frac{k}{2}} \exp\left\{-\frac{1}{2\epsilon} |G(x)|^2 \right\}. \]
So our original guess becomes
\[ \left(\frac{1}{2\pi \epsilon}\right)^\frac{k}{2} \int f(x) \exp\left\{ -\frac{1}{2\epsilon} |G(x)|^2 \right\} dx .\]
As \(\epsilon \to 0\), this integral localizes on the locus \(\{G(x) = 0\}\), as desired, but does not give the right answer! To see this, let \(u\) be a coordinate on \(M = G^{-1}(0)\) and \(v\) coordinates normal to \(M\). Then we have
\[ G(x) = G(u,v) = v^T H(u) v + o(|v|^3) \]
where \(H(x)\) is the Hessian of \(|G|^2\) at the point \(x = (u, 0)\). So the integral becomes (as \(\epsilon \to 0\))
\begin{align}
I_\epsilon &= \left(\frac{1}{2\pi \epsilon}\right)^\frac{k}{2} \int\int
f(u, v) \exp\left\{ -\frac{1}{2\epsilon} v^T H(u) v \right\} du dv \\
&= \int_M \frac{f(u)}{\sqrt{\det H(u)}} du.
\end{align}
This is not correct. We have to account for the determinant of the Hessian. Now, the Hessian is given by
\begin{align} H_{ij} &= \frac{1}{2} \frac{\partial^2 |G|^2}{\partial v^i \partial v^j} \\
&= \frac{\partial}{\partial v^i} \left( G^a \partial_j G^a \right) \\
&= \left(\partial_i G^a \partial_j G^a + G^a \partial_{ij} G^a \right) \\
&= \partial_i G^a \partial_j G^a
\end{align}
where we have used the fact that \(G = 0\) on \(x = (u, 0)\). Hence we see that
\[ \det H = (\det A)^2 \]
where \(A\) is the \(k \times k\) matrix with entries \(\partial_i G^a\). Hence
\[ \sqrt{\det H} = \det A. \]
Now there is a straightforward way to eliminate the determinant. We introduce Fermionic coordinates \(\eta^i, \theta^i\), \(i = 1, \ldots, k\). Then by Berezin integration, we have
\[ \int e^{\eta^i G^i(0, \theta^j)} d\theta d\eta = \det A. \]
So in the end, we find
\[ \int_M f(x) d\mu = \int_{\mathbb{R}^n}
f(x) \delta(G(x)) \exp\left(\eta \cdot G(x+ \theta) \right)
dx d\theta d\eta. \]
Gauge-Invariance and Gauge-Fixing
First we'll review the Fedeev-Popov method, to motivate the introduction of ghosts. Suppose we have a gauge theory involving a \(G\)-connection \(A\) and some field \(\phi\) charged under \(G\). Under a gauge transformation \(g(x)\), \(\phi\) transforms as
\[ \phi \mapsto g \cdot \phi. \]
We would like that the covariant derivative transforms in the same way, i.e.
\[ \nabla \phi \mapsto g \nabla \phi. \]
In terms of the connection 1-form \(A\), the covariant derivative is
\[ \nabla = d + A. \]
Let \(\nabla'\) denote the gauge-transformed covariant derivative, and \(\phi'\) the gauge-transformed field. Then we want
\[ \nabla' \phi' = g \nabla \phi. \]
We compute
\begin{align}
\nabla' \phi' &= (d + A')(g \phi) \\
&= dg \phi + g d\phi + A'(g\phi) \\
&= gg^{-1}dg \phi + gd\phi + g g^{-1} A' g\phi \\
&= g(d\phi + g^{-1} A' g \phi + g^{-1} dg \phi \\
&= g(d\phi + A\phi)
\end{align}
Comparing terms, we see that
\[ A = g^{-1} A' g + g^{-1} dg, \]
so upon re-arranging we have
\[ A' = g A g^{-1} - dg g^{-1} \]
This causes a problem: at critical points of the action, the Hessian of the action is degenerate in directions tangent to the gauge orbits. This means that the propagator is not well-defined, and there is no obvious way to derive the Feynman rules for perturbation theory. The solution is to take the quotient by gauge-transformations. To do this, we pick some gauge-fixing function \(G(A)\) which ought to be transverse to the orbits. Then we can restrict to the space \(G(A) = 0\), on which the Hessian of the action is non-degenerate, leading to a well-defined propagator. Formally, the path integral is
\[ Z = \int_{\{G(A) = 0\}} \exp{\frac{i}{\hbar} S[A, \phi]} \mathcal{D}A \mathcal{D}\phi \]
Formally, this suggests that the path integral should be something like
\[ Z = \int \delta(G(A)) \exp{\frac{i}{\hbar} S[A, \phi]} \mathcal{D}A \mathcal{D}\phi, \]
but this is not quite right! To understand the source of the problem, we'll first study the finite-dimensional case and then use this to solve the problem in infinite-dimensions.
The Fadeev-Popov Determinant
Suppose we are on \(\mathbb{R}^n\), and we would like to integrate a function \(f(x)\) over a submanifold \(M\) defined by \(M = G^{-1}(0)\) for some smooth function \(G: \mathbb{R}^n \to \mathbb{R}^k\). Naively, we might expect that the answer is
\[ \int_M f(x) \stackrel{?}{=} \int f(x) \delta(G(x)) dx. \]
To see why this is not correct, write the delta function as
\[ \delta(G(x)) = \frac{1}{(2\pi)^k}\int e^{ip\cdot G(x)} d^k p. \]
We can regularize this by taking the limit as \(\epsilon \to 0\) of
\[ \frac{1}{(2\pi)^k}\int \exp\left\{ip\cdot G(x) -\frac{\epsilon}{2} |p|^2\right\} d^k p \]
This integral is Gaussian, so we obtain explicitly
\[ \left(\frac{2\pi}{\epsilon} \right)^{\frac{k}{2}} \exp\left\{-\frac{1}{2\epsilon} |G(x)|^2 \right\}. \]
So our original guess becomes
\[ \left(\frac{1}{2\pi \epsilon}\right)^\frac{k}{2} \int f(x) \exp\left\{ -\frac{1}{2\epsilon} |G(x)|^2 \right\} dx .\]
As \(\epsilon \to 0\), this integral localizes on the locus \(\{G(x) = 0\}\), as desired, but does not give the right answer! To see this, let \(u\) be a coordinate on \(M = G^{-1}(0)\) and \(v\) coordinates normal to \(M\). Then we have
\[ G(x) = G(u,v) = v^T H(u) v + o(|v|^3) \]
where \(H(x)\) is the Hessian of \(|G|^2\) at the point \(x = (u, 0)\). So the integral becomes (as \(\epsilon \to 0\))
\begin{align}
I_\epsilon &= \left(\frac{1}{2\pi \epsilon}\right)^\frac{k}{2} \int\int
f(u, v) \exp\left\{ -\frac{1}{2\epsilon} v^T H(u) v \right\} du dv \\
&= \int_M \frac{f(u)}{\sqrt{\det H(u)}} du.
\end{align}
This is not correct. We have to account for the determinant of the Hessian. Now, the Hessian is given by
\begin{align} H_{ij} &= \frac{1}{2} \frac{\partial^2 |G|^2}{\partial v^i \partial v^j} \\
&= \frac{\partial}{\partial v^i} \left( G^a \partial_j G^a \right) \\
&= \left(\partial_i G^a \partial_j G^a + G^a \partial_{ij} G^a \right) \\
&= \partial_i G^a \partial_j G^a
\end{align}
where we have used the fact that \(G = 0\) on \(x = (u, 0)\). Hence we see that
\[ \det H = (\det A)^2 \]
where \(A\) is the \(k \times k\) matrix with entries \(\partial_i G^a\). Hence
\[ \sqrt{\det H} = \det A. \]
Now there is a straightforward way to eliminate the determinant. We introduce Fermionic coordinates \(\eta^i, \theta^i\), \(i = 1, \ldots, k\). Then by Berezin integration, we have
\[ \int e^{\eta^i G^i(0, \theta^j)} d\theta d\eta = \det A. \]
So in the end, we find
\[ \int_M f(x) d\mu = \int_{\mathbb{R}^n}
f(x) \delta(G(x)) \exp\left(\eta \cdot G(x+ \theta) \right)
dx d\theta d\eta. \]
Thursday, December 13, 2012
The Weyl and Wigner Transforms
Today I'd like to try to understand better how deformation quantization is related to the usual canonical quantization, and especially how the latter might be used to deduce the former, i.e., given an honest quantization (in the sense of operators), how might be reproduce the formula for the Moyal star product?
We'll fix our symplectic manifold once and for all to be \(\mathbb{R}^2\) with its standard symplectic structure, with Darboux coordinates \(x\) and \(p\). Let \(\mathcal{A}\) be the algebra of observables on \(\mathbb{R}^2\). For technical reasons, we'll restrict to those smooth functions that are polynomially bounded in the momentum coordinate (but of course the star product makes sense in general). Let \(\mathcal{D}\) be the algebra of pseudodifferential operators on \(\mathbb{R}\). We want to define a quantization map
\[ \Psi: \mathcal{A} \to \mathcal{D} \]
such that
\[ \Psi(x) = x \in \mathcal{D} \]
\[ \Psi(p) = -i\hbar \partial \]
Out of thin air, let us define
\[ \langle q| \Psi(f) |q' \rangle = \int e^{ik(q-q')} f(\frac{q+q'}{2}, k) dk \]
This is the Weyl transform. Its inverse is the Wigner transform, given by
\[ \Phi(A, q, k) = \int e^{-ikq'} \left\langle q+\frac{q'}{2} \right| A \left| q - \frac{q'}{2} \right\rangle dq' \]
Note: I am (intentionally) ignoring all factors of \(2\pi\) involved. It's not hard to work out what they are, but annoying to keep track of them in calculations, so I won't.
Theorem For suitably well-behaved \(f\), we have \( \Phi(\Psi(f)) = f\).
Proof Using the "ignore \(2\pi\)" conventions, we have the formal identities
\[ \int e^{ikx} dx = \delta(k), \ \ \int e^{ikx} dk = \delta(x). \]
The theorem is a formal result of these:
\begin{align} \Phi(\Psi(f))(q, k) &= \int e^{-ikq'} \left\langle q + \frac{q'}{2} \right| \Psi(f) \left| q - \frac{q'}{2} \right\rangle \\\
&= \int e^{-ikq'} e^{ik'q'} f(q, k) dk' dq' \\\
&= f(q,k).
\end{align}
One may easily check that \(\Psi(x) = x\) and \(Psi(k) = -i\partial\), so this certainly gives a quantization. But why is it particularly natural? To see this, let \(Q\) be the operator of multiplication by \(x\), and let \(P\) be the operator \(-i\partial\). We'd like to take \(f(q,p)\) and replace it by \(f(Q, P)\), but we can't literally substitute like this due to order ambiguity. However, we could work formally as follows:
\begin{align}
f(Q, P) &= \int \delta(Q-q) \delta(P - p) f(q,p) dq dp \\\
&= \int e^{ik(Q-q) + iq'(P-p)} f(q,p) dq dq' dp dk.
\end{align}
In this last expression, there is no order ambiguity in the argument of the exponential (since it is a sum and not a product), and furthermore the expression itself make sense since it is the exponential of a skew-adjoint operator. So let's check that this agrees with the Weyl transform. Using a special case of the Baker-Campbell-Hausdorff formula for the Heisenberg algebra, we have
\[ e^{ik(Q-q) + iq'(P-p)} = e^{ik(Q-q)} e^{iq'(P-p)} e^{-ikq'/2} \]
Let us compute the matrix element:
\begin{align}
\langle q_1 | P | q_2 \rangle &= \int \langle q_1 | p_1 \rangle
\langle p_1 | P | p_2 \rangle \langle p_2 | q_2 \rangle dp_1 dp_2 \\\
&= \int e^{iq_1p_1 - iq_2 p_2} p_2 \delta(p_2 - p_1) dp_1 dp_2 \\\
&= \int e^{i p(q_1-q_2)} p dp.
\end{align}
Hence we find that the matrix element for the exponential is
\begin{align} \langle q_1 |e^{ik(Q-q) + iq'(P-p)} | q_2 \rangle
&= e^{-ikq'/2 + ik(q_1-q)} \langle q_1 | e^{iq'(P-p)} | q_2 \rangle \\\
&= \int e^{-ikq'/2 + ik(q_1-q) -iq'p} e^{iq'p'' + ip''(q_1-q_2)} dp'' \\\
&= \delta(q' + q_1 - q_2) e^{-ikq'/2 + ik(q_1-q) -iq'p}
\end{align}
Plugging this back into the expression for \(f(Q, P)\) we find
\begin{align}
\langle q_1 | f(Q. P) | q_2 \rangle &= \int \delta(q' + q_1 - q_2) e^{-ikq'/2 + ik(q_1-q) -iq'p}
f(q,p) dq dq' dp dk \\\
&= \int e^{ ik(q_1/2 +q_2/2-q) -ip(q_1-q_2)} f(q,p) dq dp dk \\\
&= \int e^{ip(q_1-q_2)} f(\frac{q_1+q_2}{2}, p) dp,
\end{align}
which is the original expression we gave for the Weyl transform.
Out of thin air, let us define
\[ \langle q| \Psi(f) |q' \rangle = \int e^{ik(q-q')} f(\frac{q+q'}{2}, k) dk \]
This is the Weyl transform. Its inverse is the Wigner transform, given by
\[ \Phi(A, q, k) = \int e^{-ikq'} \left\langle q+\frac{q'}{2} \right| A \left| q - \frac{q'}{2} \right\rangle dq' \]
Note: I am (intentionally) ignoring all factors of \(2\pi\) involved. It's not hard to work out what they are, but annoying to keep track of them in calculations, so I won't.
Theorem For suitably well-behaved \(f\), we have \( \Phi(\Psi(f)) = f\).
Proof Using the "ignore \(2\pi\)" conventions, we have the formal identities
\[ \int e^{ikx} dx = \delta(k), \ \ \int e^{ikx} dk = \delta(x). \]
The theorem is a formal result of these:
\begin{align} \Phi(\Psi(f))(q, k) &= \int e^{-ikq'} \left\langle q + \frac{q'}{2} \right| \Psi(f) \left| q - \frac{q'}{2} \right\rangle \\\
&= \int e^{-ikq'} e^{ik'q'} f(q, k) dk' dq' \\\
&= f(q,k).
\end{align}
One may easily check that \(\Psi(x) = x\) and \(Psi(k) = -i\partial\), so this certainly gives a quantization. But why is it particularly natural? To see this, let \(Q\) be the operator of multiplication by \(x\), and let \(P\) be the operator \(-i\partial\). We'd like to take \(f(q,p)\) and replace it by \(f(Q, P)\), but we can't literally substitute like this due to order ambiguity. However, we could work formally as follows:
\begin{align}
f(Q, P) &= \int \delta(Q-q) \delta(P - p) f(q,p) dq dp \\\
&= \int e^{ik(Q-q) + iq'(P-p)} f(q,p) dq dq' dp dk.
\end{align}
In this last expression, there is no order ambiguity in the argument of the exponential (since it is a sum and not a product), and furthermore the expression itself make sense since it is the exponential of a skew-adjoint operator. So let's check that this agrees with the Weyl transform. Using a special case of the Baker-Campbell-Hausdorff formula for the Heisenberg algebra, we have
\[ e^{ik(Q-q) + iq'(P-p)} = e^{ik(Q-q)} e^{iq'(P-p)} e^{-ikq'/2} \]
Let us compute the matrix element:
\begin{align}
\langle q_1 | P | q_2 \rangle &= \int \langle q_1 | p_1 \rangle
\langle p_1 | P | p_2 \rangle \langle p_2 | q_2 \rangle dp_1 dp_2 \\\
&= \int e^{iq_1p_1 - iq_2 p_2} p_2 \delta(p_2 - p_1) dp_1 dp_2 \\\
&= \int e^{i p(q_1-q_2)} p dp.
\end{align}
Hence we find that the matrix element for the exponential is
\begin{align} \langle q_1 |e^{ik(Q-q) + iq'(P-p)} | q_2 \rangle
&= e^{-ikq'/2 + ik(q_1-q)} \langle q_1 | e^{iq'(P-p)} | q_2 \rangle \\\
&= \int e^{-ikq'/2 + ik(q_1-q) -iq'p} e^{iq'p'' + ip''(q_1-q_2)} dp'' \\\
&= \delta(q' + q_1 - q_2) e^{-ikq'/2 + ik(q_1-q) -iq'p}
\end{align}
Plugging this back into the expression for \(f(Q, P)\) we find
\begin{align}
\langle q_1 | f(Q. P) | q_2 \rangle &= \int \delta(q' + q_1 - q_2) e^{-ikq'/2 + ik(q_1-q) -iq'p}
f(q,p) dq dq' dp dk \\\
&= \int e^{ ik(q_1/2 +q_2/2-q) -ip(q_1-q_2)} f(q,p) dq dp dk \\\
&= \int e^{ip(q_1-q_2)} f(\frac{q_1+q_2}{2}, p) dp,
\end{align}
which is the original expression we gave for the Weyl transform.
Thursday, November 29, 2012
Equations of Motion and Noether's Theorem in the Functional Formalism
First, let us recall the derivation of the equations of motion and Noether's theorem in classical field theory. We have some action functional \(S[\phi]\) defined by some local Lagrangian:
\[ S[\phi] = \int L(\phi, \partial \phi) dx. \]
The classical equations of motion are just the Euler-Lagrange equations
\[ \frac{\delta S}{\delta \phi(x)} = 0
\iff \partial_\mu \left( \frac{\partial L}{\partial(\partial_\mu\phi)} \right)
= \frac{\partial L}{\partial \phi} \]
Now suppose that \(S\) is invariant under some transformation \(\phi(x) \mapsto \phi(x) + \epsilon(x) \eta(x)\), so that \(S[\phi] = S[\phi+\epsilon \eta]\). Here we treat \(\eta\) as a fixed function but \(\epsilon\) may be an arbitrary infinitesimal function. The Lagrangian is not necessarily invariant, but rather can transform with a total derivative:
\[ L(\phi+\epsilon \eta) = L(\phi)
+ \frac{\partial L}{\partial (\partial_\mu \phi)} \eta \partial_\mu \epsilon
+ \epsilon \partial_\mu f^\mu \]
For some unknown vector field \(f^\mu\) (which we could compute given any particular Lagrangian). So let's compute
\begin{align}\delta_\epsilon S &= \int \delta_\epsilon L \\\
&= \int \frac{\partial L}{\partial (\partial_\mu \phi)}
\eta \partial_\mu \epsilon + \epsilon \partial_\mu f^\mu \\\
&= \int \partial_\mu \left(f^\mu - \frac{\partial L}{\partial (\partial_\mu \phi)} \eta \right) \epsilon
\end{align}
Let us define the Noether current \(J^\mu\) by
\[ J^\mu = \frac{\partial L}{\partial (\partial_\mu \phi)} \eta - f^\mu. \]
Then the previous computation showed that
\[ \frac{\delta S}{\delta \epsilon} = -\partial_\mu J^\mu. \]
If \(\phi\) is a solution to the Euler-Lagrange equations, then the variation \(dS\) vanishes, hence we obtain:
Theorem (Noether's theorem) The Noether current is divergence free, i.e.
\[ \partial_\mu J^\mu = 0.\]
First, we derive the functional analogue of the classical equations of motion. Consider an expectation value
\[ \langle \mathcal{O(\phi)} \rangle
= \int \mathcal{O}(\phi) e^{\frac{i}{\hbar} S} \mathcal{D}\phi \]
We'll assume that \(\phi\) takes values in a vector space (or bundle). Then we can perform a change of variables \(\psi = \phi + \epsilon\), and since \(\mathcal{D}\phi = \mathcal{D}\psi\) we find that
\[ \int \mathcal{O}(\phi+\epsilon) \exp\left(\frac{i}{\hbar} S[\phi] \right) \mathcal{D}\phi \]
is independent of \(\epsilon\). Expanding to first order in \(\epsilon\), we have
\[ 0 = \int \left(\frac{\delta\mathcal{O}}{\delta \phi}
+ \frac{i \mathcal{O}}{\hbar} \frac{\delta S}{\delta \phi} \right)
\exp \left( \frac{i}{\hbar} S \right) \mathcal{D}\phi \]
So we find the quantum analogue of the equations of motion:
\[ \left\langle \frac{\delta \mathcal{O}}{\delta \phi} \right\rangle
+ \frac{i}{\hbar} \left\langle \mathcal{O} \frac{\delta S}{\delta \phi} \right\rangle = 0\]
Next, we move on to the quantum version of Noether's theorem. Suppose there is a transformation \(Q\) of the fields leaving the action invariant. Assuming the path integral measure is invariant, we obtain
\[ \left\langle QF \right \rangle + \frac{i}{\hbar} \left\langle F QS \right\rangle = 0\]
To compare with the classical result, consider \(Q\) to be the (singular) operator
\[ Q = \frac{\delta}{\delta \epsilon(x)} \]
Then by the previous calculations,
\[ Q S = -\delta_\mu J^\mu, \]
so we obtain
\[ \left\langle \frac{\delta \mathcal{O}}{\delta \epsilon(x)} \right\rangle
= \frac{i}{\hbar} \left\langle \mathcal{O} \partial_\mu J^\mu \right\rangle. \]
This is the Ward-Takahashi identity, the quantum analogue of Noether's theorem.
\[ S[\phi] = \int L(\phi, \partial \phi) dx. \]
The classical equations of motion are just the Euler-Lagrange equations
\[ \frac{\delta S}{\delta \phi(x)} = 0
\iff \partial_\mu \left( \frac{\partial L}{\partial(\partial_\mu\phi)} \right)
= \frac{\partial L}{\partial \phi} \]
Now suppose that \(S\) is invariant under some transformation \(\phi(x) \mapsto \phi(x) + \epsilon(x) \eta(x)\), so that \(S[\phi] = S[\phi+\epsilon \eta]\). Here we treat \(\eta\) as a fixed function but \(\epsilon\) may be an arbitrary infinitesimal function. The Lagrangian is not necessarily invariant, but rather can transform with a total derivative:
\[ L(\phi+\epsilon \eta) = L(\phi)
+ \frac{\partial L}{\partial (\partial_\mu \phi)} \eta \partial_\mu \epsilon
+ \epsilon \partial_\mu f^\mu \]
For some unknown vector field \(f^\mu\) (which we could compute given any particular Lagrangian). So let's compute
\begin{align}\delta_\epsilon S &= \int \delta_\epsilon L \\\
&= \int \frac{\partial L}{\partial (\partial_\mu \phi)}
\eta \partial_\mu \epsilon + \epsilon \partial_\mu f^\mu \\\
&= \int \partial_\mu \left(f^\mu - \frac{\partial L}{\partial (\partial_\mu \phi)} \eta \right) \epsilon
\end{align}
Let us define the Noether current \(J^\mu\) by
\[ J^\mu = \frac{\partial L}{\partial (\partial_\mu \phi)} \eta - f^\mu. \]
Then the previous computation showed that
\[ \frac{\delta S}{\delta \epsilon} = -\partial_\mu J^\mu. \]
If \(\phi\) is a solution to the Euler-Lagrange equations, then the variation \(dS\) vanishes, hence we obtain:
Theorem (Noether's theorem) The Noether current is divergence free, i.e.
\[ \partial_\mu J^\mu = 0.\]
Functional Version
First, we derive the functional analogue of the classical equations of motion. Consider an expectation value
\[ \langle \mathcal{O(\phi)} \rangle
= \int \mathcal{O}(\phi) e^{\frac{i}{\hbar} S} \mathcal{D}\phi \]
We'll assume that \(\phi\) takes values in a vector space (or bundle). Then we can perform a change of variables \(\psi = \phi + \epsilon\), and since \(\mathcal{D}\phi = \mathcal{D}\psi\) we find that
\[ \int \mathcal{O}(\phi+\epsilon) \exp\left(\frac{i}{\hbar} S[\phi] \right) \mathcal{D}\phi \]
is independent of \(\epsilon\). Expanding to first order in \(\epsilon\), we have
\[ 0 = \int \left(\frac{\delta\mathcal{O}}{\delta \phi}
+ \frac{i \mathcal{O}}{\hbar} \frac{\delta S}{\delta \phi} \right)
\exp \left( \frac{i}{\hbar} S \right) \mathcal{D}\phi \]
So we find the quantum analogue of the equations of motion:
\[ \left\langle \frac{\delta \mathcal{O}}{\delta \phi} \right\rangle
+ \frac{i}{\hbar} \left\langle \mathcal{O} \frac{\delta S}{\delta \phi} \right\rangle = 0\]
Next, we move on to the quantum version of Noether's theorem. Suppose there is a transformation \(Q\) of the fields leaving the action invariant. Assuming the path integral measure is invariant, we obtain
\[ \left\langle QF \right \rangle + \frac{i}{\hbar} \left\langle F QS \right\rangle = 0\]
To compare with the classical result, consider \(Q\) to be the (singular) operator
\[ Q = \frac{\delta}{\delta \epsilon(x)} \]
Then by the previous calculations,
\[ Q S = -\delta_\mu J^\mu, \]
so we obtain
\[ \left\langle \frac{\delta \mathcal{O}}{\delta \epsilon(x)} \right\rangle
= \frac{i}{\hbar} \left\langle \mathcal{O} \partial_\mu J^\mu \right\rangle. \]
This is the Ward-Takahashi identity, the quantum analogue of Noether's theorem.
Saturday, November 24, 2012
The Moyal Product
Today I want to understand the Moyal product, as we will need to understand it in order to construct quantizations of symplectic quotients. (More precisely, to incorporate stability conditions.)
Let \(A\) be the algebra of polynomial functions on \(T^\ast \mathbb{C}^n\). This algebra has a natural Poisson bracket, given by
\[ \{p_i, x_j\} = \delta_{ij}. \]
We would like to define a new associative product \(\ast\) on \(A((\hbar))\) satisfying:
\[ f \ast g = m \circ B(f \otimes g). \]
Now, condition (2) tells us that \(B(0) = 1\) and that
\[ \left. \frac{dB}{d\hbar} \right|_{\hbar=0} = \frac{\Pi}{2} \]
So
\[ B = 1 + \frac{\hbar \Pi}{2} + O(\hbar^2) \]
It is natural to guess that \(B\) should be built out of powers of \(\Pi\), and a natural guess is
\[ B = \exp(\frac{\hbar \Pi}{2}), \]
which certainly reproduces the first two terms of our expansion. Let's see that this choice actually works, i.e. defines an associative \(\ast\)-product. Let \(m: A \otimes A \to A\) be the multiplication, and
\(m_{12}, m_{23}: A \otimes A \otimes A \to A \otimes A\), \(m_{123}: A \otimes A \otimes A \to A\) the induced multiplication maps. Then
\begin{align}
f \ast (g \ast h) &= m \circ(B( f \otimes m \circ B(g \otimes h) ) ) \\\
&= m \circ B( m_{23} \circ (1 \otimes B)(f \otimes g \otimes h) ) \\\
&= m_{123} (B \otimes 1)(1 \otimes B)(f \otimes g \otimes h)
\end{align}
On the other hand, we have
\begin{align}
(f \ast g) \ast h) &= m \circ(B( m \circ B(f \otimes g) \otimes h) ) ) \\\
&= m \circ B( m_{12} \circ (B \otimes 1)(f \otimes g \otimes h) ) \\\
&= m_{123} (1 \otimes B)(B \otimes 1)(f \otimes g \otimes h)
\end{align}
Hence, associativity is the condition
\[ m_{123} \circ [1\otimes B, B \otimes 1] = 0. \]
On \(A \otimes A \otimes A\), write \(\partial_i^1\) for the partial derivative acting on the first factor, \(\partial_i^2\) on the second, etc. Then
\[ 1 \otimes B = \sum_n \frac {\hbar^n}{2^n n!}
\Pi^{i_1 j_1} \cdots \Pi^{i_n j_n} \partial^2_{i_1} \partial^3_{j_1} \cdots
\partial^2_{i_n} \partial^3_{j_n} \]
and similarly for \(B \otimes 1\). So we have
\begin{align}
m_{123} (B\otimes 1)(1 \otimes B) &= \sum_n \sum_{k=0}^n \frac {\hbar^n}{2^n k! (n-k)!}
\Pi^{k_1 l_1} \cdots \Pi^{k_k l_k} \partial_{k_1} \partial_{l_1} \cdots
\partial_{k_k} \partial_{l_k} \\\
& \ \times \Pi^{i_1 j_1} \cdots \Pi^{i_{n-k} j_{n-k}} \partial_{i_1} \partial_{j_1} \cdots
\partial_{i_{n-k}} \partial_{j_{n-k}} \\\
&= m_{123}(1 \otimes B)(B \otimes 1)
\end{align}
Hence we obtain an associative \(\ast\)-product. This is called Moyal product.
Now suppose that \(U\) is a (Zariski) open subset of \(X = T^\ast \mathbb{C}^n\). Then the star product induces a well-defined map
\[ \ast: O_X(U)((\hbar)) \otimes_\mathbb{C} O_X(U)((\hbar)) \to O_X(U)((\hbar)) \]
In this way we obtain a sheaf \(\mathcal{D}\) of \(O_X\) modules with a non-commutative \(\ast\)-product defined as above.
Define a \(\mathbb{C}^\ast\) action on \(T^\ast \mathbb{C}^n\) by acting on \(x_i\) and \(p_i\) with weight 1. Extend this to an action on \(\mathcal{D}\) by acting on \(hbar\) with weight -1.
Proposition: The algebra \(C^\ast\)-invariant global sections of \(\mathcal{D}\) is naturally identified with the algebra of differential operators on \(\mathbb{C}^n\).
Proof: The \(\mathbb{C}^\ast\)-invariant global sections are generated by \(\hbar^{-1} x_i\) and \(\hbar^{-1} p_i\). So define a map \(\Gamma(\mathcal{D})^{\mathbb{C}^\ast} \to \mathbb{D}\) by
\[ \hbar^{-1} x_i \mapsto x_i \]
\[ \hbar^{-1} p_i \mapsto \partial_i \]
From the definition of the star product, it is clear that this is an algebra map, and that it is both injective and surjective.
Let \(A\) be the algebra of polynomial functions on \(T^\ast \mathbb{C}^n\). This algebra has a natural Poisson bracket, given by
\[ \{p_i, x_j\} = \delta_{ij}. \]
We would like to define a new associative product \(\ast\) on \(A((\hbar))\) satisfying:
- \(f \ast g = fg + O(\hbar) \)
- \(f \ast g - g \ast f = \hbar \{f, g\} + O(\hbar^2)\)
- \(1 \ast f = f \ast 1 = f\)
- \((f \ast g)^\ast = -g^\ast \ast f^\ast\)
In the last line, the map \((\cdot)^\ast\) takes \(x_i \mapsto x_i\) and \(p_i \mapsto -p_i\). To figure out what this new product should be, let's take \(f,g \in A\) and expand \(f \ast g\) in power series:
\[ f \ast g = \sum_{n=0}^\infty c_n(f,g) \hbar^n \]
Now, equations (1) and (2) will be satisfied by taking \(c_0(f,g) = fg\) and \(c_1(f,g) = \{f,g\}/2\). Let \(\sigma\) be the Poisson bivector defining the Poisson bracket. This defines a differential operator \(\Pi\) on \(A \otimes A\) by
\[ \Pi = \sigma^{ij} (\partial_i \otimes \partial_j) \]
Let \(B = \sum_{n=0}^\infty B_n \hbar^n\) and write the product as
\[ f \ast g = m \circ B(f \otimes g). \]
Now, condition (2) tells us that \(B(0) = 1\) and that
\[ \left. \frac{dB}{d\hbar} \right|_{\hbar=0} = \frac{\Pi}{2} \]
So
\[ B = 1 + \frac{\hbar \Pi}{2} + O(\hbar^2) \]
It is natural to guess that \(B\) should be built out of powers of \(\Pi\), and a natural guess is
\[ B = \exp(\frac{\hbar \Pi}{2}), \]
which certainly reproduces the first two terms of our expansion. Let's see that this choice actually works, i.e. defines an associative \(\ast\)-product. Let \(m: A \otimes A \to A\) be the multiplication, and
\(m_{12}, m_{23}: A \otimes A \otimes A \to A \otimes A\), \(m_{123}: A \otimes A \otimes A \to A\) the induced multiplication maps. Then
\begin{align}
f \ast (g \ast h) &= m \circ(B( f \otimes m \circ B(g \otimes h) ) ) \\\
&= m \circ B( m_{23} \circ (1 \otimes B)(f \otimes g \otimes h) ) \\\
&= m_{123} (B \otimes 1)(1 \otimes B)(f \otimes g \otimes h)
\end{align}
On the other hand, we have
\begin{align}
(f \ast g) \ast h) &= m \circ(B( m \circ B(f \otimes g) \otimes h) ) ) \\\
&= m \circ B( m_{12} \circ (B \otimes 1)(f \otimes g \otimes h) ) \\\
&= m_{123} (1 \otimes B)(B \otimes 1)(f \otimes g \otimes h)
\end{align}
Hence, associativity is the condition
\[ m_{123} \circ [1\otimes B, B \otimes 1] = 0. \]
On \(A \otimes A \otimes A\), write \(\partial_i^1\) for the partial derivative acting on the first factor, \(\partial_i^2\) on the second, etc. Then
\[ 1 \otimes B = \sum_n \frac {\hbar^n}{2^n n!}
\Pi^{i_1 j_1} \cdots \Pi^{i_n j_n} \partial^2_{i_1} \partial^3_{j_1} \cdots
\partial^2_{i_n} \partial^3_{j_n} \]
and similarly for \(B \otimes 1\). So we have
\begin{align}
m_{123} (B\otimes 1)(1 \otimes B) &= \sum_n \sum_{k=0}^n \frac {\hbar^n}{2^n k! (n-k)!}
\Pi^{k_1 l_1} \cdots \Pi^{k_k l_k} \partial_{k_1} \partial_{l_1} \cdots
\partial_{k_k} \partial_{l_k} \\\
& \ \times \Pi^{i_1 j_1} \cdots \Pi^{i_{n-k} j_{n-k}} \partial_{i_1} \partial_{j_1} \cdots
\partial_{i_{n-k}} \partial_{j_{n-k}} \\\
&= m_{123}(1 \otimes B)(B \otimes 1)
\end{align}
Hence we obtain an associative \(\ast\)-product. This is called Moyal product.
Sheafifying the Construction
Now suppose that \(U\) is a (Zariski) open subset of \(X = T^\ast \mathbb{C}^n\). Then the star product induces a well-defined map
\[ \ast: O_X(U)((\hbar)) \otimes_\mathbb{C} O_X(U)((\hbar)) \to O_X(U)((\hbar)) \]
In this way we obtain a sheaf \(\mathcal{D}\) of \(O_X\) modules with a non-commutative \(\ast\)-product defined as above.
Define a \(\mathbb{C}^\ast\) action on \(T^\ast \mathbb{C}^n\) by acting on \(x_i\) and \(p_i\) with weight 1. Extend this to an action on \(\mathcal{D}\) by acting on \(hbar\) with weight -1.
Proposition: The algebra \(C^\ast\)-invariant global sections of \(\mathcal{D}\) is naturally identified with the algebra of differential operators on \(\mathbb{C}^n\).
Proof: The \(\mathbb{C}^\ast\)-invariant global sections are generated by \(\hbar^{-1} x_i\) and \(\hbar^{-1} p_i\). So define a map \(\Gamma(\mathcal{D})^{\mathbb{C}^\ast} \to \mathbb{D}\) by
\[ \hbar^{-1} x_i \mapsto x_i \]
\[ \hbar^{-1} p_i \mapsto \partial_i \]
From the definition of the star product, it is clear that this is an algebra map, and that it is both injective and surjective.
Subscribe to:
Posts (Atom)