Mathematics for Machine Learning
Notes are grouped by conceptual cluster and ordered from foundational to advanced within each cluster. Each cluster feeds into the next.
I. Algebraic Structures
The algebraic scaffolding that everything else is built on.
- Vectors — The basic objects: what they are and how they behave
- Closure — When an operation keeps you inside the set
- Groups — Sets with a well-behaved operation (closure, inverse, identity)
- Vector Spaces — Groups of vectors with scalar multiplication
- Linear Independence — Non-redundancy; each vector adds new information
- Basis — A minimal, spanning set; the coordinate system of a space
II. Matrices & Linear Systems
How equations are represented and solved using matrices.
- Matrix — Rectangular arrays; the language of linear algebra
- System of Linear Equations — The central problem: find x such that Ax = b
- Row View of a System of Linear Equations — Each equation as a hyperplane
- Column View of a System of Linear Equations — Finding the right linear combination of columns
- Matrix View of a System of Linear Equations — A as an operator transforming x into b
- Gaussian Elimination — The algorithmic workhorse for solving linear systems
- Inverse and Transpose — Undoing a matrix transformation; flipping rows and columns
- Rank — The effective dimensionality of a matrix’s output
- Image and Kernel — The output space (image) and collapse space (kernel); rank-nullity theorem
- Determinant — Volume scaling factor; zero iff non-invertible
- Trace — Sum of diagonal elements; sum of eigenvalues
III. Linear Mappings & Transformations
How matrices encode functions between vector spaces.
- Linear Mappings — Functions preserving addition and scalar multiplication
- Matrices as Linear Transformation — The matrix as a geometric operator on space
- Matrix Representation of Linear Mappings — Encoding a linear map as a matrix w.r.t. a basis
- Views of Matrix Multiplication — Four perspectives: dot product, column, row, outer product
- Rotations — Linear transformations that preserve distances and angles
IV. Geometry & Inner Products
Measuring length, distance, angles, and projections.
- Inner Products — Bilinear maps that generalize the dot product
- Norms — Length functions derived from inner products
- Lengths and Distances — Inner products induce norms; Cauchy-Schwarz inequality
- Angles and Orthogonality — Angles between vectors; orthogonal and orthonormal bases
- Orthogonal Projections — Projecting a vector onto a subspace; least-squares geometry
- Gram-Schmidt Orthogonalization — Constructing an orthonormal basis from any basis
- Inner Product of Functions — Extending inner products to function spaces via integration
- Matrices as Inner Products — SPD matrices define custom geometries; Mahalanobis distance
- Symmetric, Positive Definite Matrices — The matrices that define valid inner products
V. Matrix Decompositions
Factoring matrices to reveal structure and enable computation.
- Matrix decomposition landscape — Overview and taxonomy of decomposition methods
- Characteristic Polynomial — det(A − λI) = 0; isolates eigenvalues as roots
- Eigenvalues and eigenvectors — Directions scaled but not rotated; spectral structure
- Eigendecomposition and Diagonalization — A = PDP⁻¹; powers, dynamics, PCA
- Cholesky Decomposition — A = LLᵀ for SPD matrices; efficient system solving
- Singular Value Decomposition — A = UΣVᵀ; the universal factorization for any matrix
VI. Calculus
Derivatives generalized to multiple variables and vector-valued functions.
- Univariate Calculus Refresher — Derivative as a limit; Taylor series; differentiation rules
- Partial Derivative — Differentiating with respect to one variable at a time; the gradient
- Gradients of Vector-Valued Functions — The Jacobian matrix; generalizing gradient to f: Rⁿ → Rᵐ
- The Jacobian Determinant — Local volume scaling factor for nonlinear transformations
- Hessian Matrix — Matrix of second-order partials; curvature; identifying critical points
VII. Probability Theory
The mathematical framework for reasoning under uncertainty.
- The Landscape of Probability Theory — Mindmap overview of the full probability landscape
- The Probability Space — The triplet (Ω, A, P); Kolmogorov axioms
- Conditional Probability — P(A|B); independence; chain rule; law of total probability
- Random Variables — RVs as functions Ω → ℝ; PMF, PDF, CDF
- Expectation and Variance — E[X], Var(X); linearity of expectation; bias-variance
- Common Probability Distributions — Bernoulli, Gaussian, Categorical, Poisson, Beta and more
- Law of Large Numbers and Central Limit Theorem — Sample mean converges to E[X]; sums of i.i.d. RVs become Gaussian
- Covariance Matrix — Cov(X,Y); the Σ matrix as an SPD matrix; Mahalanobis distance
- MLE and MAP — MLE as log-likelihood maximisation; MAP as MLE + prior; connection to loss functions and regularisation
- Bayes Theorem — Posterior ∝ likelihood × prior; the engine of Bayesian ML
VIII. Bridges
Notes that connect two clusters. Each lives at the intersection of two bodies of knowledge.
Linear Algebra ↔ Probability:
- Geometry of Random Variables — Variance = squared length; correlation = cosine angle; RVs form a Hilbert space
- Orthogonal Projections as Conditional Expectation — E[Y|X] is the orthogonal projection of Y onto the subspace of functions of X; why MSE loss targets the conditional mean
Calculus ↔ Probability:
- Transformations of Probability Densities — How a PDF changes under a nonlinear map; the Jacobian determinant as a density scaling factor
Cross-Cluster Connections
Key conceptual bridges between clusters: