Index

This index points to the section, definition, or theorem where each topic is introduced, rather than to a page number, so that it remains equally useful in the PDF and web editions of this book.

A

  • Adjusted Rand Index, Definition 6.3
  • Association rules (out of scope for this book), 1
  • Autoencoders, 5.9

B

  • Bayesian Information Criterion (BIC), 6.4
  • Bottleneck (autoencoder), 5.9

C

  • Centering matrix, 1
  • Charts and atlases, 5.3
  • Classical (metric) scaling, 4.4
  • Cluster validation, 6.6
  • Clustering, overview, Chapter 6
    • Center-based, 6.1
    • Hierarchical, 6.3
    • Kernel \(k\)-means, 6.2
    • Model-based (GMM), 6.4
    • Spectral, 6.5
  • Cophenetic correlation, Definition 6.1
  • Covariance constraints (GMM), 6.4

D

  • Dendrograms, 6.3
  • Determinant, 1
  • Dimension reduction, overview, 1

E

  • Eckart-Young-Mirsky Theorem, Theorem 4.3
  • Eigenmaps, see Laplacian eigenmap, Hessian eigenmaps
  • Elbow plot, 6.1
  • Expectation-Maximization (EM) algorithm, 6.4

F

  • Feature map, 5.1

G

  • Gap statistic, 6.1
  • Gaussian Mixture Models, 6.4
  • Geodesic distance, 5.4
  • Graph Laplacian, 5.6

H

  • Hessian eigenmaps (HLLE), 5.7
  • Hierarchical clustering, 6.3

I

  • Intrinsic dimension, 5.3
  • Isometry, Definition 5.6
  • ISOMAP, 5.4

K

  • \(k\)-center, 6.1
  • \(k\)-means, 6.1
    • Kernel \(k\)-means, 6.2
    • Lloyd’s algorithm, 6.1
  • \(k\)-medoids, 6.1
  • Kernel PCA, 5.2
  • Kernel trick, 5.1

L

  • Laplacian eigenmap, 5.6
  • Linkage (single, complete, average, Ward), 6.3
  • Loadings (PCA), 4.1
  • Locally Linear Embeddings (LLE), 5.5

M

  • Manifold, Definition 5.1
  • Manifold distance, Definition 5.5
  • Manifold hypothesis, 3.1
  • Manifold learning, overview, Chapter 5
  • Mercer’s condition, Theorem 5.1
  • MNIST dataset, 1
  • Multidimensional Scaling (MDS), 4.4

N

  • NCI60 tumor microarray dataset, 6.6
  • Neighborhood graph, 5.4
  • Nonnegative Matrix Factorization (NMF), 4.3
  • Normalized Mutual Information, Definition 6.4

P

  • Principal Component Analysis (PCA), 4.1
    • First PCA loading, Definition 4.1
    • Optimality of PCA, Theorem 4.2
    • PCA and eigendecomposition, Theorem 4.1

R

  • Rand Index, Definition 6.2
  • Recommendation systems, 4.2
  • Reinforcement learning, 1

S

  • Scree plot, 4.1
  • Self-supervised learning, 1
  • Semi-supervised learning, 1
  • Silhouette score, 6.1.5.3
  • Singular Value Decomposition (SVD), 4.2
  • Spectral clustering, 6.5
  • Supervised learning, 1

T

  • Tangent space, 5.3

U

  • UMAP, 5.8
  • Unifying view (manifold learning as kernel PCA), 5.11
  • Unsupervised learning, definition, 1

W

  • Ward linkage, 6.3
[1]
Deng, L. (2012). The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine 29 141–2.
[2]
Hastie, T., Tibshirani, R. and Friedman, J. (2001). The elements of statistical learning. Springer New York Inc., New York, NY, USA.
[3]
Izenman, A. J. (2008). Modern multivariate statistical techniques: Regression, classification, and manifold learning. Springer Publishing Company, Incorporated.
[4]
Trefethen, L. N. and Bau, D. (1997). Numerical linear algebra. SIAM.
[5]
Strang, G. (2006). Linear algebra and its applications. Thomson, Brooks/Cole, Belmont, CA.
[6]
Campadelli, P., Casiraghi, E., Ceruti, C. and Rozza, A. (2015). Intrinsic dimension estimation: Relevant techniques and a benchmark framework. Mathematical Problems in Engineering 2015 759567.
[7]
Spivak, M. D. (1979). A comprehensive introduction to differential geometry. Publish or Perish, Inc.
[8]
Lee, J. M. (2019). Introduction to riemannian manifolds (second edition). Springer Nature.
[9]
S., K. P. F. R. (1901). LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 2 559–72.
[10]
Hotelling, H. (1933). Analysis of a complex of statistical variables into principal components. Journal of educational psychology 24 417.
[11]
Donoho, D., Gavish, M. and Romanov, E. (2023). ScreeNOT: Exact MSE-optimal singular value thresholding in correlated noise. The Annals of Statistics 51 122–48.
[12]
Cattell, R. B. (1966). The scree test for the number of factors. Multivariate Behavioral Research 1 245–76.
[13]
Lee, D. and Seung, H. S. (2000). Algorithms for non-negative matrix factorization. In Advances in neural information processing systems vol 13, (T. Leen, T. Dietterich and V. Tresp, ed). MIT Press.
[14]
Liu, W., Tang, A., Ye, D. and Ji, Z. (2008). Nonnegative singular value decomposition for microarray data analysis of spermatogenesis. In 2008 international conference on information technology and applications in biomedicine pp 225–8.
[15]
Brunet, J.-P., Tamayo, P., Golub, T. R. and Mesirov, J. P. (2004). Metagenes and molecular pattern discovery using matrix factorization. Proceedings of the National Academy of Sciences 101 4164–9.
[16]
Cutler, A. and Breiman, L. (1994). Archetypal analysis. Technometrics 36 338–47.
[17]
Lin, C.-H., Ma, W.-K., Li, W.-C., Chi, C.-Y. and Ambikapathi, A. (2014). Identifiability of the simplex volume minimization criterion for blind hyperspectral unmixing: The no pure-pixel case. IEEE Transactions on Geoscience and Remote Sensing 53.
[18]
Fu, X., Ma, W.-K., Huang, K. and Sidiropoulos, N. D. (2015). Blind separation of quasi-stationary sources: Exploiting convex geometry in covariance domain. IEEE Transactions on Signal Processing 63 1–1.
[19]
Févotte, C., Bertin, N. and Durrieu, J.-L. (2009). Nonnegative matrix factorization with the itakura-saito divergence: With application to music analysis. Neural Computation 21 793–830.
[20]
Cox, T. F. and Cox, M. A. A. (2000). Multidimensional scaling, second edition. CRC Press.
[21]
Alam, M. A. and Fukumizu, K. (2014). Hyperparameter selection in kernel principal component analysis. Journal of Computer Science 10 1139–50.
[22]
Mika, S., Schölkopf, B., Smola, A., Müller, K.-R., Scholz, M. and Rätsch, G. (1998). Kernel PCA and de-noising in feature spaces. In Advances in neural information processing systems vol 11, (M. Kearns, S. Solla and D. Cohn, ed). MIT Press.
[23]
Little, A., Lee, J., Jung, Y.-M. and Maggioni, M. (2009). Estimation of intrinsic dimensionality of samples from noisy low-dimensional manifolds in high dimensions with multiscale SVD. In IEEE Workshop on Statistical Signal Processing Proceedings pp 85–8.
[24]
Tenenbaum, J. B., Silva, V. de and Langford, J. C. (2000). A global geometric framework for nonlinear dimensionality reduction. Science 290 2319–23.
[25]
Bernstein, M., Silva, V., Langford, J. and Tenenbaum, J. (2001). Graph approximations to geodesics on embedded manifolds.
[26]
Roweis, S. T. and Saul, L. K. (2000). Nonlinear dimensionality reduction by locally linear embedding. Science 290 2323–6.
[27]
Chen, J. and Liu, Y. (2011). Locally linear embedding: A survey. Artif. Intell. Rev. 36 29–48.
[28]
[29]
Anon. (2019). Locally linear embedding with additive noise. Pattern Recognition Letters 123 47–52.
[30]
Chang, H. and Yeung, D.-Y. (2006). Robust locally linear embedding. Pattern Recognition 39 1053–65.
[31]
Belkin, M. and Niyogi, P. (2001). Laplacian eigenmaps and spectral techniques for embedding and clustering. In Advances in neural information processing systems vol 14, (T. Dietterich, S. Becker and Z. Ghahramani, ed). MIT Press.
[32]
Donoho, D. L. and Grimes, C. (2003). Hessian eigenmaps: Locally linear embedding techniques for high-dimensional data. Proceedings of the National Academy of Sciences 100 5591–6.
[33]
McInnes, L., Healy, J. and Melville, J. (2018). UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426.
[34]
Hinton, G. E. and Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science 313 504–7.
[35]
Kingma, D. P. and Ba, J. (2015). Adam: A method for stochastic optimization. In International conference on learning representations (ICLR).
[36]
Vincent, P., Larochelle, H., Bengio, Y. and Manzagol, P.-A. (2008). Extracting and composing robust features with denoising autoencoders. In ICML ’08 pp 1096–103. Association for Computing Machinery.
[37]
Kingma, D. P. and Welling, M. (2013). Auto-encoding variational bayes. CoRR abs/1312.6114.
[38]
Ham, J., Lee, D. D., Mika, S. and Schölkopf, B. (2004). A kernel view of the dimensionality reduction of manifolds. In Proceedings of the twenty-first international conference on machine learning (ICML) p 47. ACM.
[39]
Dhillon, I. S., Guan, Y. and Kulis, B. (2004). Kernel k-means: Spectral clustering and normalized cuts. In Proceedings of the tenth ACM SIGKDD international conference on knowledge discovery and data mining pp 551–6. ACM.
[40]
Roux, M. (2018). A comparative study of divisive and agglomerative hierarchical clustering algorithms. Journal of Classification 35 345–66.
[41]
Szekely, G. J., Rizzo, M. L., et al. (2005). Hierarchical clustering via joint between-within distances: Extending ward’s minimum variance method. Journal of classification 22 151–84.
[42]
Everitt, B. S. (2001). Cluster analysis. Arnold ; Oxford University Press, London : New York.
[43]
Mojena, R. (1977). Hierarchical grouping methods and stopping rules: an evaluation*. The Computer Journal 20 359–63.
[44]
Dempster, A. P., Laird, N. M. and Rubin, D. B. (1977). Maximum Likelihood from Incomplete Data Via the EM Algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology 39 1–22.
[45]
Banfield, J. D. and Raftery, A. E. (1993). Model-based gaussian and non-gaussian clustering. Biometrics 49 803–21.
[46]
Karlis, D. and Santourian, A. (2009). Model-based clustering with non-elliptically contoured distributions. Statistics and Computing 19 73–83.
[47]
O’Hagan, A., Murphy, T. B., Gormley, I. C., McNicholas, P. D. and Karlis, D. (2016). Clustering with the multivariate normal inverse Gaussian distribution. Computational Statistics & Data Analysis 93 18–30.
[48]
Dang, U. J., Gallaugher, M. P. B., Browne, R. P. and McNicholas, P. D. (2023). Model-Based Clustering and Classification Using Mixtures of Multivariate Skewed Power Exponential Distributions. Journal of Classification 40 145–67.
[49]
Dombowsky, A. and Dunson, D. B. (2025). Bayesian Clustering via Fusing of Localized Densities. Journal of the American Statistical Association.