Mathematics for Machine Learning

Notes are grouped by conceptual cluster and ordered from foundational to advanced within each cluster. Each cluster feeds into the next.


I. Algebraic Structures

The algebraic scaffolding that everything else is built on.

  1. Vectors — The basic objects: what they are and how they behave
  2. Closure — When an operation keeps you inside the set
  3. Groups — Sets with a well-behaved operation (closure, inverse, identity)
  4. Vector Spaces — Groups of vectors with scalar multiplication
  5. Linear Independence — Non-redundancy; each vector adds new information
  6. Basis — A minimal, spanning set; the coordinate system of a space

II. Matrices & Linear Systems

How equations are represented and solved using matrices.

  1. Matrix — Rectangular arrays; the language of linear algebra
  2. System of Linear Equations — The central problem: find x such that Ax = b
  3. Gaussian Elimination — The algorithmic workhorse for solving linear systems
  4. Inverse and Transpose — Undoing a matrix transformation; flipping rows and columns
  5. Rank — The effective dimensionality of a matrix’s output
  6. Image and Kernel — The output space (image) and collapse space (kernel); rank-nullity theorem
  7. Determinant — Volume scaling factor; zero iff non-invertible
  8. Trace — Sum of diagonal elements; sum of eigenvalues

III. Linear Mappings & Transformations

How matrices encode functions between vector spaces.

  1. Linear Mappings — Functions preserving addition and scalar multiplication
  2. Matrices as Linear Transformation — The matrix as a geometric operator on space
  3. Matrix Representation of Linear Mappings — Encoding a linear map as a matrix w.r.t. a basis
  4. Views of Matrix Multiplication — Four perspectives: dot product, column, row, outer product
  5. Rotations — Linear transformations that preserve distances and angles

IV. Geometry & Inner Products

Measuring length, distance, angles, and projections.

  1. Inner Products — Bilinear maps that generalize the dot product
  2. Norms — Length functions derived from inner products
  3. Lengths and Distances — Inner products induce norms; Cauchy-Schwarz inequality
  4. Angles and Orthogonality — Angles between vectors; orthogonal and orthonormal bases
  5. Orthogonal Projections — Projecting a vector onto a subspace; least-squares geometry
  6. Gram-Schmidt Orthogonalization — Constructing an orthonormal basis from any basis
  7. Inner Product of Functions — Extending inner products to function spaces via integration
  8. Matrices as Inner Products — SPD matrices define custom geometries; Mahalanobis distance
  9. Symmetric, Positive Definite Matrices — The matrices that define valid inner products

V. Matrix Decompositions

Factoring matrices to reveal structure and enable computation.

  1. Matrix decomposition landscape — Overview and taxonomy of decomposition methods
  2. Characteristic Polynomial — det(A − λI) = 0; isolates eigenvalues as roots
  3. Eigenvalues and eigenvectors — Directions scaled but not rotated; spectral structure
  4. Eigendecomposition and Diagonalization — A = PDP⁻¹; powers, dynamics, PCA
  5. Cholesky Decomposition — A = LLᵀ for SPD matrices; efficient system solving
  6. Singular Value Decomposition — A = UΣVᵀ; the universal factorization for any matrix

VI. Calculus

Derivatives generalized to multiple variables and vector-valued functions.

  1. Univariate Calculus Refresher — Derivative as a limit; Taylor series; differentiation rules
  2. Partial Derivative — Differentiating with respect to one variable at a time; the gradient
  3. Gradients of Vector-Valued Functions — The Jacobian matrix; generalizing gradient to f: Rⁿ → Rᵐ
  4. The Jacobian Determinant — Local volume scaling factor for nonlinear transformations
  5. Hessian Matrix — Matrix of second-order partials; curvature; identifying critical points

VII. Probability Theory

The mathematical framework for reasoning under uncertainty.

  1. The Landscape of Probability Theory — Mindmap overview of the full probability landscape
  2. The Probability Space — The triplet (Ω, A, P); Kolmogorov axioms
  3. Conditional Probability — P(A|B); independence; chain rule; law of total probability
  4. Random Variables — RVs as functions Ω → ℝ; PMF, PDF, CDF
  5. Expectation and Variance — E[X], Var(X); linearity of expectation; bias-variance
  6. Common Probability Distributions — Bernoulli, Gaussian, Categorical, Poisson, Beta and more
  7. Law of Large Numbers and Central Limit Theorem — Sample mean converges to E[X]; sums of i.i.d. RVs become Gaussian
  8. Covariance Matrix — Cov(X,Y); the Σ matrix as an SPD matrix; Mahalanobis distance
  9. MLE and MAP — MLE as log-likelihood maximisation; MAP as MLE + prior; connection to loss functions and regularisation
  10. Bayes Theorem — Posterior ∝ likelihood × prior; the engine of Bayesian ML

VIII. Bridges

Notes that connect two clusters. Each lives at the intersection of two bodies of knowledge.

Linear Algebra ↔ Probability:

  1. Geometry of Random Variables — Variance = squared length; correlation = cosine angle; RVs form a Hilbert space
  2. Orthogonal Projections as Conditional Expectation — E[Y|X] is the orthogonal projection of Y onto the subspace of functions of X; why MSE loss targets the conditional mean

Calculus ↔ Probability:

  1. Transformations of Probability Densities — How a PDF changes under a nonlinear map; the Jacobian determinant as a density scaling factor

Cross-Cluster Connections

Key conceptual bridges between clusters:

FromToWhy
Symmetric, Positive Definite MatricesInner ProductsSPD matrices define valid inner products
DeterminantCharacteristic Polynomialdet(A − λI) defines the char. poly
Eigenvalues and eigenvectorsEigendecomposition and DiagonalizationEigenvalues/vectors are the ingredients
RankImage and Kernelrank(A) = dim(Im(Φ)); rank-nullity theorem
Orthogonal ProjectionsGram-Schmidt OrthogonalizationG-S iteratively applies projections
BasisMatrix Representation of Linear MappingsA basis is needed to write down a transformation matrix
Angles and OrthogonalityRotationsRotations preserve angles and distances
The Probability SpaceBayes TheoremConditional probability from the probability space
Inner ProductsGeometry of Random VariablesRVs with finite variance form an inner product space
Covariance MatrixSymmetric, Positive Definite MatricesThe covariance matrix is always symmetric positive semi-definite
Orthogonal ProjectionsOrthogonal Projections as Conditional ExpectationConditional expectation IS an orthogonal projection in L²
The Jacobian DeterminantTransformations of Probability DensitiesThe Jacobian determinant scales density under a change of variables
Expectation and VarianceLaw of Large Numbers and Central Limit TheoremLLN and CLT are theorems about limits of sample means