Preparations for the upcoming SJTU internship

In this chapter I am writing about my preparations for the upcoming Summer Research Internship at Shanghai Jiao Tong University. The internship will cover: Equivariance, SE(3) Transformations, Lie Groups, Flow Matching (!) for Motion Planning, Brownian Bridge, Context-Aware Transport, and Diffusion Models for Motion Planning. I will start from the fundamental math covered, and will proceed to more advanced concepts.

The First Chapter consists of Lie Groups, SE(3) Transformations, and Equivariance.

Lie Groups

A group is a set \(G\) with a binary operation \(\cdot: G \times G \to G\) that combines any two elements of the set \((a, b) \in G \times G\) to produce a third element \(a \cdot b\) within the same set \(G\) and the following conditions must hold: the operation is associative \(\forall a, b, c \in G, \, (a \cdot b) \cdot c = a \cdot (b \cdot c)\), it has an identity element \(\exists e \in G\) such that \(\forall a \in G, \, e \cdot a = a \cdot e = a\), and every element of the set has an inverse element \(\forall a \in G, \, \exists a^{-1} \in G\) such that \(a \cdot a^{-1} = a^{-1} \cdot a = e\). For example, the integers \(\mathbb{Z}\) with the addition operation \(+\) form a group \((\mathbb{Z}, +)\).

A Lie group (pronounced li, i think because the creator had his name pronounced like that) is a group \(G\) that is also a differentiable manifold \(\mathcal{M}\), such that group multiplication \(\mu: G \times G \to G\) and taking inverses \(\iota: G \to G\) are both differentiable \(C^\infty\) maps.

A key example of what differs a Lie Group from any other is the type/amount of operations we can do to our set. Imagine rotating an object such that after the operation it stays the same, defining its symmetry group. For a square this happens every 90 degrees, forming the discrete cyclic group \(\mathbb{Z}_4\). But for a circle there are an infinite, uncountable amount of angles \(\theta \in [0, 2\pi)\) such that the circle stays the same, forming the continuous 1-dimensional Lie group \(SO(2)\).

Discrete vs. Continuous Rotation

0°

A differentiable Manifold is a manifold that is differentiable. A Manifold is any topological space where we can zoom into each point and it resembles an Euclidean Space. Example the number 8 is not a manifold because the intersection in the middle, however much you zoom into it, will never resemble an Euclidean space. (In 2 dimensions the Euclidean space is a line. any other point in 8 will have a neighbordhood which resembles a line, but the intersection will never.) For a space to be a 1D topological manifold, every single point must have a small neighborhood that can be stretched into a simple, flat open line segment

Summarizing, for a line-like shape, it needs to be able to be topologically transformed into a line, and a 3-dimensional shape needs to be able to be topologically transformed into a plane.

1: How can a group be differentiable? How can a Lie group be a differentiable manifold?

To understand this, you need to understand that you don't take the derivative of the group elements themselves; but you take the derivative of the group actions (multiplication and inversion). A group is differentiable because you can execute the group operations infinitesimally.

Formally: For a group to be differentiable, its two core operations—multiplication \((g_1, g_2) \mapsto g_1 \cdot g_2\) and inversion \(g \mapsto g^{-1}\)—must behave like smooth mathematical functions. If you perturb the inputs by an infinitesimal step (\(dt\)), the output responds smoothly (\(dy\)), allowing you to compute a standard derivative:

$$\frac{d}{dt}(g_1(t) \cdot g_2(t))$$

To state that a Lie group \(G\) is a group that is also a differentiable manifold means that \(G\) simultaneously satisfies both algebraic and differential-topological axioms. These two independent structures are bound together by something called a compatibility condition: the fundamental group operations must be smooth maps between manifolds.

2. The Smooth Compatibility Axioms

Let \(G\) be a smooth (\(C^\infty\)) manifold and a group under the binary operation \(\cdot\). We define two foundational mappings:

For \(G\) to be classified as a Lie group, both \(\mu\) and \(\iota\) must be smooth maps. This implies that if we evaluate these operations within local coordinate charts, the resulting transition functions must possess continuous partial derivatives of all orders.

3: The Concrete Example: Invertible Matrices

The best way to see how a group becomes a smooth manifold is to look at the most famous Lie group of all: the General Linear Group, denoted as \(GL(n, \mathbb{R})\). This is simply the set of all \(n \times n\) real matrices that have an inverse.

Mathematically, we write this set as:

$$GL(n, \mathbb{R}) = \{ A \in \mathbb{M}_{n \times n}(\mathbb{R}) \mid \det(A) \neq 0 \}$$

4. Why is it a Manifold?

Think about a standard \(2 \times 2\) matrix. It has 4 numbers inside it. If we unfold those 4 numbers into a flat line, it looks exactly like a coordinate point in 4D space (\(\mathbb{R}^4\)). In general, any \(n \times n\) matrix can be thought of as a single point sitting inside a flat \(n^2\)-dimensional space (\(\mathbb{R}^{n^2}\)).

The condition for a matrix to belong to our group is that its determinant cannot equal zero (\(\det(A) \neq 0\)). Because the determinant is just a smooth polynomial equation, the places where \(\det(A) = 0\) form thin, sharp "walls" cutting through our space.

By throwing out the matrices on those walls, we are left with open, continuous chambers of invertible matrices. Because any point inside an open chamber has plenty of breathing room to move around in all directions without hitting a wall, the space locally behaves exactly like flat space. This makes \(GL(n, \mathbb{R})\) an open submanifold of \(\mathbb{R}^{n^2}\).

5. Why is Multiplication Smooth?

To be a Lie group, combining two elements must be a differentiable process. Let's see what happens when we multiply two matrices, \(A\) and \(B\), to get a output matrix \(C = AB\). We calculate every individual slot \(c_{ij}\) using the standard row-by-column dot product formula:

$$c_{ij} = a_{i1}b_{1j} + a_{i2}b_{2j} + \dots + a_{in}b_{nj} = \sum_{k=1}^{n} a_{ik}b_{kj}$$

Look closely at what this operation is actually doing to our coordinates. It is just multiplying pairs of numbers and adding them together. In calculus, we know that basic polynomials are infinitely differentiable—there are no fractions with variables in the denominator, no sharp absolute values, and no square roots. Because the output coordinates are just polynomials of the input coordinates, the multiplication map is automatically smooth (\(C^\infty\)).

6. Why is Inversion Smooth?

The final test is finding the inverse. If we take a matrix \(A\) and calculate \(A^{-1}\), we can use Cramer's Rule from linear algebra to write out the formula explicitly:

$$A^{-1} = \frac{1}{\det(A)} \text{adj}(A)$$

This formula tells us that every single number in our inverted matrix is a fraction: the top part (the adjugate matrix) is made of simple polynomials, and the bottom part is the determinant polynomial.

In calculus, a rational function (a fraction of polynomials) is perfectly differentiable everywhere except where the denominator is zero. But remember our group's golden rule: we already threw away every single matrix where \(\det(A) = 0\)! Because the denominator is guaranteed to never be zero anywhere on our manifold, the inversion map is completely smooth (\(C^\infty\)).

7: The Big Picture

Because matrix multiplication and matrix inversion are made up of clean, fraction-safe algebra, we can use calculus on them without running into undefined steps or sharp breaks. This compatibility is exactly what allows us to look at the tangent space at the identity matrix, build the flat Lie Algebra, and use the exponential map to smoothly optimize transformations.

SE(3) Transformations

A \(SE(3)\) Transformation is often called a rigid-body motion. The \(SE(3)\) is the set of all rigid-body transformations that can be done to an object. Rigid Body = Rotation + Translation. Nothing else. $$T(\mathbf{x}) = R\mathbf{x} + \mathbf{t} \quad \text{where} \quad R \in SO(3), \mathbf{t} \in \mathbb{R}^3$$ \(SE(3)\) means Special Euclidean Group. This Group is a subset of the Euclidean Groups. The Euclidean group is the group of (Euclidean) isometries of an Euclidean space \(E^n\); that is, the transformations of that space that preserve the Euclidean distance between any two points (also called Euclidean transformations). $$||T(\mathbf{x}) - T(\mathbf{y})||_2 = ||\mathbf{x} - \mathbf{y}||_2$$ The key difference is that Euclidean Groups, along Rotation and Translation, contain also Reflection (turning objects inside out). $$\det(R) = \pm 1 \quad \text{(Euclidean Group)} \quad \text{vs.} \quad \det(R) = 1 \quad \text{(Special Euclidean Group)}$$

\(SE(3)\) is a smooth manifold (a Lie group). One can't just add two \(SE(3)\) matrices together because the result won't be a valid transformation anymore. $$T_1, T_2 \in SE(3) \implies T_1 + T_2 \notin SE(3)$$ This makes optimization (in calculus or machine learning) really hard. To get around this, One can map \(SE(3)\) to its tangent space at the identity, known as the Lie Algebra, \(\mathfrak{se}(3)\). $$T = \exp(\mathbf{v}^\wedge) \quad \text{where} \quad \mathbf{v} \in \mathfrak{se}(3)$$

1: Talking about \(\mathfrak{se}(3)\)

If you are training a neural network or running an optimization algorithm, you constantly need to update your position: New Position = Old Position + Step. But: if you take two points on a globe and draw a straight line between them using normal addition, that line cuts through the inside of the globe. You are no longer on the surface.

BUT if we use a tangent space (tangent plane to the point we are at right now), since it is a flat 2D plane (standard Euclidean vector space), we can draw straight lines on it, add vectors together, and do standard calculus without breaking any mathematical rules

Lie algebra is the tangent space at the identity.

2: Mapping between \(SE(3)\) and \(\mathfrak{se}(3)\)

So, how do we actually "map" \(SE(3)\) to its tangent space and back? We use two special functions:

(The hat operator ^ is explained later)

3: The Homogeneous Transformation Matrix

While I defined \(SE(3)\) conceptually as rotation plus translation, in practice (Python/C++) I need to represent elements of \(SE(3)\) using \(4 \times 4\) homogeneous transformation matrices.

If I have a rotation matrix \(R \in SO(3)\) and a translation vector \(\mathbf{t} \in \mathbb{R}^3\), the corresponding \(SE(3)\) matrix is constructed as:

$$T = \begin{bmatrix} R & \mathbf{t} \\ \mathbf{0}^T & 1 \end{bmatrix} = \begin{bmatrix} r_{11} & r_{12} & r_{13} & t_x \\ r_{21} & r_{22} & r_{23} & t_y \\ r_{31} & r_{32} & r_{33} & t_z \\ 0 & 0 & 0 & 1 \end{bmatrix}$$

This \(4 \times 4\) structure is mathematically brilliant because it turns the non-linear operation of rotation and translation (\(R\mathbf{x} + \mathbf{t}\)) into a single, linear matrix multiplication. To transform a 3D point \(\mathbf{x} = [x, y, z]^T\), I append a \(1\) to make it homogeneous (\(\tilde{\mathbf{x}} = [x, y, z, 1]^T\)) and simply multiply: \(\tilde{\mathbf{y}} = T\tilde{\mathbf{x}}\).

4: \(\mathfrak{se}(3)\) (Twists)

In the previous section, I noted that \(\mathfrak{se}(3)\) is the tangent space at the identity, but what does an element of \(\mathfrak{se}(3)\) actually look like?

An element of the Lie algebra \(\mathfrak{se}(3)\) is represented by a 6-dimensional vector called a twist, usually denoted as \(\boldsymbol{\xi} \in \mathbb{R}^6\). It consists of linear velocity \(\mathbf{v}\) and angular velocity \(\boldsymbol{\omega}\):

$$\boldsymbol{\xi} = \begin{bmatrix} \mathbf{v} \\ \boldsymbol{\omega} \end{bmatrix} \in \mathbb{R}^6$$

To map this 6D vector to the \(4 \times 4\) matrix space (so we can use the exponential map \(\exp(\boldsymbol{\xi}^\wedge)\)), we use the hat operator (\(^\wedge\)). The hat operator takes the twist vector and constructs a \(4 \times 4\) matrix where the angular velocity becomes a skew-symmetric matrix (\(\boldsymbol{\omega}^\wedge\)):

$$\boldsymbol{\xi}^\wedge = \begin{bmatrix} \boldsymbol{\omega}^\wedge & \mathbf{v} \\ \mathbf{0}^T & 0 \end{bmatrix} = \begin{bmatrix} 0 & -\omega_z & \omega_y & v_x \\ \omega_z & 0 & -\omega_x & v_y \\ -\omega_y & \omega_x & 0 & v_z \\ 0 & 0 & 0 & 0 \end{bmatrix}$$

This matrix \(\boldsymbol{\xi}^\wedge\) is the actual element of the Lie algebra. The exponential map then "integrates" this velocity over time \(t=1\) to give you the final \(SE(3)\) pose matrix \(T\).

Equivariance

When applying transformations (like our \(SE(3)\) rotations and translations) to the input of a neural network or a mathematical function, we generally look for one of two behaviors: Invariance or Equivariance.

1. Invariance

A function is invariant if transforming the input does nothing to the output. An example to this is classification. Mathematically, if \(f\) is our network and \(T\) is our transformation:

$$f(T(\mathbf{x})) = f(\mathbf{x})$$

2. Equivariance

A function is equivariant if transforming the input causes the output to transform in the exact same mathematical way. Imagine a robotic motion planning network predicting the exact 3D coordinate to grasp that coffee mug. If you translate the mug by vector \(\mathbf{t}\), the predicted grasp coordinate must also translate by \(\mathbf{t}\).

$$f(T(\mathbf{x})) = T(f(\mathbf{x}))$$

3: Why is this crucial for Motion Planning?

Standard neural networks are not inherently equivariant to \(SE(3)\). If I train a standard network to grasp a mug sitting upright, and then hand it a mug lying on its side, it will fail. The most intuitive solution was data augmentation, forcing the network to memorize objects at every possible angle.

By designing network architectures that are mathematically guaranteed to be \(SE(3)\)-Equivariant, I can build the laws of 3D physics directly into the model's structure. If the network learns how to compute a motion plan for an object in one orientation, it immediately and perfectly generalizes to every other \(SE(3)\) pose without needing extra training data.

← Back to blog