II.2 Unsupervised
learning
1
LE THI HOAI AN
University of Lorraine
2
Outline
ο΄ Clustering
Hard clustering, Fuzzy clustering,
Hierarchical clustering,
Weighted clustering,
Clustering by maximizing Modularity,
Self Organizing Map (SOM)
ο΄ Principal Component Analysis (PCA)
ο΄ Singular Value Decomposition (SVD)
Clustering
4
Clustering
A technique to find similar
groups in data clusters
General Objective:
Given: A dataset of π points in π-dimensional real space
Problem: Extract hidden distinct properties by clustering the
dataset into π clusters
⇒ Choose π members as centroid (or “median”) and assign
each member to its closest centroid.
5
Clustering
Clustering is a fundamental problem in unsupervised
learning which has many applications in various domains.
In recent years, there has been significant interest in
developing clustering algorithms to the massive data sets.
Two main approaches have been studied for clustering:
ο΄ the statistical and machine learning based on learning
mixture models
ο΄ the mathematical programming approach that considers
clustering as an optimization problem.
6
Clustering
The general term “clustering” covers many different
types of problems.
All consist of subdividing a data set into groups of
similar elements,
But there are many measures of similarity, many ways
of measuring, and various concepts of subdivision.
7
Clustering
8
Clustering
ο΄ Clustering rows:
grouping similar objects
ο΄ Clustering columns:
grouping similar variables across
samples
ο΄ Bi-clusteing/Two-way clustering:
grouping aobjects that are similar
across a subset of variables
9
Clustering
Two essential components of cluster analysis:
ο΄ Distance measure:
A notion of distance or similarity of two objects:
When are two objects close to each other?
ο΄ Cluster algorithm:
A procedure to minimize distance of objects
within groups and/or maximize distance between
groups
10
Examples of distance measure
ο΄ Euclidean distance measure
average
difference
across
coordinates
ο΄ Manhattan distance measures
average
difference
across
coordinates in a robust way
Manhattan distance: the red, yellow, and blue paths all
have the shortest length of 12
Euclidean distance: the green line has length 6 2 =
8.49, and is the unique shortest path.
11
Clustering algorithms
ο΄ Popular algorithms for clustering
ο§ Hard clustering
ο§ Hierarchical clustering
ο§ K-mean
ο§ SOMs (Self-Organizing Maps)
ο§ Autoclass, mixture models
ο΄ Hierarchical clustering allows the choice of the
dissimilarity matrix
ο΄ K-mean and SOMs take orginal data directly as input.
Attributes are assumed to live in Eulcidean space.
Hard Clustering
12
Minimum Sum-of-Squares Clustering
ο΄ An instance of the partitional clustering problem consists
of a data set π΄ β {π1 , … , ππ } of π points in βπ , a
measured distance, and an integer π.
ο΄ Goal: Choose π members π₯ β (β = 1, … π) (in π΄ and/or
βπ ) as ”centroid” (or ”median”) and assign each member
of A to its closest centroid.
ο΄ The assignment distance of a point π ∈ π΄ is the distance
from π to the centroid to which it is assigned,
ο΄ The objective function, which is to be minimized, is the
sum of assignment distances.
13
Hard Clustering
ο΄ If the centroids are not necessarily in π΄, then the
problem can be formulated as a unconstrained
optimization problem.
ο΄ In the contrary case, we are faced with a discrete
optimization problem.
ο΄ In both cases, different objective functions
corresponding to the distance metric being considered
are possible.
ο΄ Two models widely studied in the literature are the
cases where the points come from a real space βπ and
the assignment distance of a point is defined as the
squared Euclidean distance (2-norm) and/or the 1norm
15
π-mean algorithm
Given a random initial set of k means, the algorithm
proceeds by alternating between two steps:
ο΄Assignment step: Assign each observation to the
cluster whose mean yields the least within-cluster
sum of squares (i.e. the "nearest" mean in
Euclidean distance)
ο΄Update step: Calculate the new means to be the
centroids of the observations in the new clusters.
19
Clustering: π-means model
If the squared Euclidean distance is used and the
centroids are not necessarily in π¨,
then the corresponding optimization problem can be
expressed as (|| ⋅ || denotes the Euclidean norm)
Bilevel model
π
min ΰ· min
π=1
β=1,…,π
2
β
π
π₯ −π
: π₯ β ∈ βπ , β = 1, … , π
20
Kernel MSSC
Consider the Gaussian kernel function π
: βπ × βπ → β
and the corresponding mapping π: βπ → π» (a Hilbert
space) such that
π₯−π¦ 2
π π₯ , π(π¦) = π
π₯, π¦ = exp −
2π 2
π>0
We have the expression
π π₯
β
π
− π(π )
2
β
π₯ −π
= 2 − 2exp −
2π 2
π 2
The optimization model of the Gaussian kernel MSSC
π
π₯ β − ππ
min ΰ· min −2exp −
β=1,…,π
2π 2
π=1
2
: π₯ β ∈ βπ , β = 1, … , π
21
Clustering:
π’π§πππ ππ« π©π«π¨π π«ππ¦π¦π’π§π π¦π¨πππ₯
Let π = π’βπ ∈ βπ×π with β = 1, … , π and π = 1, … , π be
the matrix defined by
1 if ππ belongs to cluster βπ‘β,
π’βπ β α
0 otherwise.
Then a straightforward mixed integer formulation of MSSC
π
π
β
ΰ· ΰ· π’βπ π₯ − π
π 2
:
π=1 β=1
min
π’βπ ∈ 0,1 , π = 1, … , π, β = 1, … , π
π
ΰ· π’βπ = 1, π = 1, … , π
β=1
22
π-median problem: bilevel model
In the π-median problem, the centers must be among the
given data set, i.e. we are to choose π members πβ ∈ π¨ as
centers and assign each member of π΄ to its closest center.
Formulation as a bilevel model
π
ΰ· min
min
π=1
π
β=1,…,π
ΰ· min
β=1
π=1,…,π
2
β
π
π₯ −π
:
π₯ β − ππ = 0
23
π-median problem
In the π-median problem, the centers must be among the
given data set, i.e. we are to choose π members πβ ∈ π¨ as
centers and assign each member of π΄ to its closest center.
Formulation as a linear 0-1 programming:
For π₯ ∈ π΄, define 0-1 variable π(π₯) by
1 if π₯ is a center,
π π₯ = α
0 otherwise.
For each pair π₯ and π¦ of π΄, define the variable π π₯, π¦ :
1 if π₯ is a center and π¦ is assigned to π₯,
π π₯, π¦ = α
0 otherwise.
24
π-median problem
The π-median (π − ππ) takes the form:
min ΰ· π π₯, π¦ π(π₯, π¦)
π₯≠π¦∈π΄
s. t.
ΰ· π π₯, π¦ ≥ 1 , ∀π¦ ∈ π΄
π₯∈π΄
π π₯, π¦ ≤ π π₯ ∀π₯, π¦ ∈ π΄
(π-MP) is
NP-hard
ΰ·π π₯ ≤ π,
π₯∈π΄
π π₯ ∈ 0,1 , ∀π₯ ∈ π΄
π π₯, π¦ ∈ 0,1 , ∀π₯, π¦ ∈ π΄
First constraint : all points π¦ of π΄ are assigned to a center.
Second one: if π¦ is assigned to π₯ then π₯ must be a center.
π π₯, π¦ : the assigned cost (“distance” between π₯ and π¦).
25
Fuzzy clustering
26
Fuzzy Clustering
ο΄ In fuzzy clustering, objects are not classified as belonging
to one and only one cluster, but instead, they all possess a
degree of membership with each of the clusters.
ο΄ In real applications, there is very often no sharp boundary
between clusters so that fuzzy clustering is often better
suited for the data.
ο΄ This work in included in the mathematical programming
approach and concerns with the fuzzy clustering. We
consider the Fuzzy C-Means (FCM) clustering model that is
undoubtedly a most widely used fuzzy clustering
technique. It was originally introduced by Bezdek in 1981
as a fuzzification of π-Means model of hard clustering.
27
Fuzzy Clustering
ο΄ Let π βΆ= {π₯1 , π₯2 , … , π₯π } denote π objects to be
partitioned into π (2 ≤ π ≤ π) homogeneous clusters
πΆ1 , πΆ2 , … , πΆπ where π₯π ∈ βπ (π = 1, … , π) represents
multispectral (features) data.
ο΄ Consider the matrix π = π’π,π π×π called the fuzzy
partition matrix in which each element π’π,π indicates
the membership degree of each object π₯π in the
cluster πΆπ (the probability that π₯π belongs to the
cluster πΆπ ).
28
Fuzzy Clustering
ο΄ The FCM technique is based on optimizing the
objective function:
where π is the π × π -matrix whose πth row is π£π ∈
βπ the conter of cluster πΆπ . π ≥ 1: the fuzziness
index of membership of each datum.
29
Fuzzy Clustering
ο΄The mathematical model of FCM is
30
Image
segmentation by fuzzy clustering
Medical image with the Gaussian noise and the results of segmentation (c=3)
31
Hierarchical clustering
32
Hierarchical clustering
ο΄ Multilevel hierarchical clustering consists of grouping data
objects into a hierarchy of clusters.
ο΄ It has a long history and has many important applications
in various domains, since many kinds of data, including
observational data collected in the human and biological
sciences, have a hierarchical, nested, or clustered
structure.
ο΄ Hierarchical clustering algorithms are useful to determine
hierarchical multicast trees for Gird computing using in eScience, e-Medicine or e-Commerce, Multimedia
conferencing, Large-scale dissemination of timely
information, etc.
33
Hierarchical clustering
ο΄ A hierarchical clustering of a set of objects can be
described as a tree, in which the leaves are precisely the
objects to be clustered.
ο΄ A hierarchical clustering scheme produces a sequence of
clusterings in which each clustering is nested into the next
clustering in the sequence.
ο΄ Standard existing methods for Multilevel hierarchical
clustering are often based upon nonhierarchical clustering
algorithms coupled with several iterative control strategies
to repeatedly modify an initial clustering (reordering, and
reclustering) in search of a better one.
34
Hierarchical clustering
ο΄ While mathematical programming is widely used for
nonhierarchical clustering problems there exist a few
optimization models and techniques for multilevel hierarchical
clustering ones.
ο΄ Problem statement
ο§ Given a set π of π objects π βΆ = {π₯π ∈ βπ : π = 1, … , π}, a
measured distance, and an integer π.
ο§ We are to choose π + 1 members in π΄, one as the total center
(the root of the tree) and others as centers of π disjoint clusters,
and assign other members of π΄ to their closest center.
ο§ The total center is defined as the closest object to all centers (in
the sense that the sum of distances between it and all centers is
the smallest).
35
Hierarchical clustering
ο΄ A practical application widely studied in the networking
community that is the network topology identification based on
end-to-end measurements.
ο΄ Consider a communication network, in which a sender node
transmits information packets to a set of receiver nodes. The
receivers are, in this case, the usual ”objects” to be clustered.
ο΄ Assume that the routes from the sender to the receivers are
fixed.
ο΄ The physical network topology is essentially a graph, where each
node corresponds to a physical device (e.g., router, switch,
terminal, etc...) and the links correspond to the connections
between them.
ο΄ Knowledge of the network topology is essential for tasks like
monitoring and provisioning a network.
36
Hierarchical clustering
The first optimization model:
37
38
Hierarchical clustering
We consider a new mathematical program in which all the
centers are variables. Denote ππ π = 1, … , π the center of
clusters in the second level and ππ+π the total center. The
third optimization model:
39
Construction of a multicast
communication network
ο΄ The identification problem of the topology of a network:
using clustering algorithms.
ο΄ Construction of a multicast communication network
• Objective: minimize the total cost of communication
• Groupware systems, Video conference, video streaming
of international events, P2P applications to share data or
processing, etc.
40
Construction of a multicast
communication network
ο΄ A hierarchical multicast tree contains a source (the total
center) with several levels of hierarchy
ο΄ A node is connected to a higher-level node (except the
source) and its lower-level nodes (if any)
44
Weighted Clustering
45
Weighted clustering
In many applications such as e-commerce applications,
computational biology, text classification, image analysis,
etc. datasets are large volume and contain a large
number of features.
Three categories of Features: relevant, redundant and
irrelevant features.
ο΄ Relevant features are essential for classification
process,
ο΄ Redundant features add no new information to the
classifier (i.e., information already carried by other
features) while
ο΄ Irrelevant features do not carry any useful information.
46
Weighted clustering
ο΄ Feature selection deals with irrelevant or redundant
features.
Goal: select a subset of features that minimize
redundancy while preserving or improving the
classification rate of algorithm.
ο΄ Feature weighting: an extension of feature selection.
In feature weighting, each feature is assigned a
continuous value, named a weight, in the interval [0,1].
Relevant features correspond to a high weight value,
whereas a weight value close to zero represent
irrelevant features.
Objective: improve the quality of classification
algorithm, but not to reduce the number of features.
47
Weighted clustering
Problem statement
Given π βΆ = {π₯π ∈ βπ : π = 1, … , π} and an integer π.
The dissimilarity measure πΏππΉ between π£π and π₯π is now
defined by π weighted features:
48
Weighted clustering
Bilevel formulation of MSSC using weighted features
49
50
51
52
Clustering by Maximizing
Modularity
53
Clustering by Maximizing
Modularity
ο΄ Community structure: one important network features.
The whole network: composed of densely connected subnetworks, with only sparser connections between them. Such
sub-networks are called communities or modules.
ο΄ Detection of communities is of significant practical
importance as it allows to analyze networks at a megascopic
scale:
• identification of related web pages in the WWW,
• uncovering of communities in social networks,
• decomposition of metabolic networks in functional
modules.
54
Clustering by Maximizing
Modularity
ο΄ Detecting communities ⇒ searching for the structure
that maximizes the number of intra-community edges
while minimizing the number of inter-community edges.
ο΄ The modularity (Q measure): characterize the existence
of community structure in a network.
ο΄ The modularity Q of a particular partition is defined as
the number of edges inside clusters, minus the expected
number of such edges if the graph were random
conditioned on its degree distribution.
55
56
57
58
59
Self Organizing Map (SOM)
60
Self Organizing Map (SOM)
ο΄ The Self-Organizing Map (SOM), introduced by
Kohonen in 1982, and its variants are a popular
artificial neural network approaches in unsupervised
learning.
ο΄ The principal goal of an SOM is to transform an
incoming signal pattern of arbitrary dimension into a
low (usually one or two)-dimensional discrete map,
and to perform this transformation adaptively in a
topologically ordered fashion.
61
Self Organizing Map (SOM)
ο΄ The SOM is an elegant way to interpret complex data and
an excellent tool for data visualization and exploratory
cluster analysis. SOM is primarily a data-driven
dimensionality reduction and data compression method.
ο΄ In data mining and machine learning, SOM is a widely
used tool in the exploratory phase for various technique
such as data clustering, data classification and graph
mining.
62
63
64
Principal component analysis
(PCA)
65
Principal component analysis (PCA)
66
Principal component analysis (PCA)
ο΄ The principal components of a collection of points are a
sequence of direction vectors, where the ith vector is the
direction of a line that best fits the data while being
orthogonal to the first i-1 vectors.
ο΄ PCA is the process of computing the principal components
and using them to perform a change of basis on the data,
sometimes using only the first few principal components
and ignoring the rest.
ο΄ Use an orthogonal transformation: convert a set of
observations of possibly correlated variables ⇒ a set of
values of linearly uncorrelated variables.
ο΄ The number of principal components ≤ that of original
variables.
67
Objectives of PCA
ο΄ Reduce dimensionality (pre-processing for other methods)
ο΄ Choose the most useful (informative) variables
ο΄ Compress the data
ο΄ Visualize multidimensional data
• To identity groups of objects
• To identity outliers
68
Example: visualization
ο΄Given 53 blood and urine samples (features) from 65
people. How can we visualize the measurements?
550
500
450
400
350
300
250
200
150
100
50
0 50
Tri-variate (3 features)
4
M-EPI
C-LDH
Bi-variate (2 features)
3
2
1
0
600
400
150 250 350 450
C-Triglycerides
C-LDH 200 0 0
500
400
300
200
100
C-Triglycerides
How can we visualize the other variables???
… difficult to see in 4 or higher dimensional spaces...
69
ο΄ Is there a representation better than the coordinate axes?
ο΄ Is it really necessary to show all the 53 dimensions?
ο΄ How could we find the smallest subspace of the 53-D
space that keeps the most information about the original
data?
ο΄ A solution: Principal Component Analysis
70
PCA Algorithm
Input: π × π data matrix π, with one row vector π₯π per
data point
Step 1: π ← subtract mean π₯ from each row vector π₯π in π
Step 2: Σ ← covariance matrix of π
Step 3: Find eigenvectors and eigenvalues of Σ
Step 4: PC’s ← the eigenvectors with largest eigenvalues
Step 5: Derive the new data
PCA Example –STEP 1
DATA:
x
y
2.5
2.4
0.5
0.7
2.2
2.9
1.9
2.2
3.1
3.0
2.3
2.7
2
1.6
1
1.1
1.5
1.6
1.1
0.9
Original data
71
PCA Example –STEP 1
• Subtract the mean from each of the data
dimensions.
All the π₯ values have π₯ subtracted and π¦ values have
π¦ subtracted from them. This produces a data set
whose mean is zero.
Subtracting the mean makes variance and covariance
calculation easier by simplifying their equations.
The variance and co-variance values are not affected
by the mean value.
72
PCA Example –STEP 1
DATA:
π₯
π¦
2.5
2.4
0.5
0.7
2.2
2.9
1.9
2.2
3.1
3.0
2.3
2.7
2
1.6
1
1.1
1.5
1.6
1.1
0.9
π₯π – mean(π₯)
π¦π – mean(π¦)
ZERO MEAN DATA:
π₯
π¦
.69
.49
-1.31
-1.21
.39
.99
.09
.29
1.29
1.09
.49
.79
.19
-.31
-.81
-.81
-.31
-.31
-.71
-1.01
73
PCA Example –STEP 2
• Calculate the covariance matrix of zero mean data
cov =
cov(π₯, π₯)
cov(π¦, π₯)
.616555556
=
.615444444
where
cov(π₯, π¦)
cov(π¦, π¦)
.615444444
.716555556
π
1
cov π₯, π¦ =
ΰ·(π₯π − mean(π₯))(π¦π − mean(π¦))
π−1
π=1
74
PCA Example –STEP 3
• Calculate the eigenvectors and eigenvalues of the
covariance matrix
ο§ Find eigenvalues π = π1 , π2 π of matrix cov:
det(cov-ππ I) = 0
.0490833989
π =
1.28402771
ο§ Find eigenvectors π£ = (π£1 , π£2 ) of cov:
cov*π£π =ππ ∗ π£π
−.735178656
π£=
.677873399
−.677873399
−.716555556
75
PCA Example –STEP 4
• order eigenvectors found from eigenvectors by
eigenvalue, highest to lowest.
-.677873399 -.735178656
-.735178656 .677873399
• choose the principle components in order of
significance.
Here, the first principle component :
- .677873399
- .735178656
76
PCA Example –STEP 5
• Deriving the new data
FinalData = RowFeatureVector x RowZeroMeanData
RowFeatureVector is the matrix with the eigenvectors
in the columns transposed so that the eigenvectors
are now in the rows, with the most significant
eigenvector at the top
RowZeroMeanData is the mean-adjusted data
transposed, ie. the data items are in each column,
with each row holding a separate dimension.
77
PCA Example –STEP 5
x
-.827970186
1.77758033
-.992197494
-.274210416
-1.67580142
-.912949103
.0991094375
1.14457216
.438046137
1.22382056
The data after transforming using only the most significant eigenvector
78
79
PCA and Optimization
ο΄ π : π × π data matrix with rows π₯π ;
ο΄ π : dimension of a subspace into which points are
projected (π ≤ π);
ο΄ π : π × π matrix of scores;
ο΄ π : π × π matrix comprised of the first π columns π£π ,
π = 1, … , π of the rotation matrix
ο΄ We need to find π, π½ such that
πΏπ» = π½ππ»
under the constraint π½π» π½ = π° where π° is the identity
matrix.
80
PCA and Optimization
ο΄ Given a data matrix π × π πΏ and π ≤ π
ο΄ Goal is to minimize the sum of squared distances of points to
their projections in a π-dimensional subspace.
π
min ΰ· π₯π − ππ§π 22
π,π
s. t. π π π = πΌ
π=1
where πΌ is the identity matrix.
• Score π§π : projection of π₯π in the π-dimensional subspace,
• Quantity π¦π = ππ§π : projected point in terms of the original
coordinates ⇒ π : π × π matrix of such projected points;
•Columns of π define a basis for the subspace.
• An optimal solution is to set π to be the loadings for the
first π PCs and π = ππ
81
Alternative optimization problem
π
min ΰ· π₯π − π¦π 22
π
π=1
s. t. ππππ π ≤ π
• In this formulation, we do not decompose the projected
points to find the fitted subspace explicitly.
• Rather, the sum of squared distances of points π₯π to their
projections in terms of the original coordinates π¦π is
minimized while restricting the dimension of the subspace
containing the projections to a value at most q.
• Setting π = πππ π where the columns of π are the
loadings of first π PCs provides an optimal solution to this
problem.
Singular Value Decomposition
(SVD)
83
Singular Value Decomposition (SVD)
ο΄ Singular value decomposition (SVD) is a
factorization of a real or complex matrix into
simpler meaningful pieces, used to detect groupings
in data
ο΄ It is the generalization of the eigen decomposition
of a positive semidefinite normal matrix (for
example, a symmetric matrix with positive
eigenvalues) to any π × π matrix via an extension
of polar decomposition.
ο΄ It has many useful applications in signal processing
and statistics.
84
Singular Value Decomposition (SVD)
For an nο΄ m matrix A of rank r there exists a factorization
(Singular Value Decomposition = SVD) as follows:
A ο½ ULV
πο΄π
πο΄π
T
πο΄π
The columns of U are orthogonal eigenvectors of AAT.
The columns of V are orthogonal eigenvectors of ATA.
Eigenvalues π1 … ππ of AAT are the eigenvalues of ATA.
ππ =
ππ
πΏ = diag π1 . . . ππ
Singular value
85
Let
SVD: Example
ο© 2 2οΉ
Aο½οͺ
οΊ
ο
1
1
ο«
ο»
ο©1 0 οΉ ο© 2 2
οͺ0 1 οΊ οͺ
ο«
ο» οͺο« 0
Thus π = π = 2. Its SVD is
0 οΉ ο© 1 2 1 2οΉ
οΊοͺ
οΊ
2 οΊο» οͺο« ο1 2 1 2 οΊο»
Typically, the singular values arranged in
decreasing order.
86
SVD and Optimization
ο΄ Given a data matrix π × π π΄
ο΄ GOAL: find π½, π³, πΌ such that
π¨ = πΌπ³π½π»
ο΄ In PCA, under the constraint π½π» π½ = π° where π° is the
identity matrix, we need to find π, π½ such that
πΏπ» = π½ππ»
• Note that based on SVD we can rewrite the objective
function π₯π − ππ§π 22 in PCA as ππ − ππΏπ’π 22 .
• Replace β2 -norm by π, π
norm: π΄ π,π
=
tr ππ΄π
π΄π .
• Adjusting π and π
: control sparsity and smoothness.
87
SVD and Optimization
The Generalized least squares Matrix Decomposition
optimization problem that generalizes SVD:
π
min ΰ· ππ − ππΏπ’π 2π,π
π=1
s. t.
π π π
π = πΌ,
π π ππ = πΌ,
diag πΏ ≥ 0
where πΌ is the appropriately-sized identity matrix,
π΄ π,π
=
tr ππ΄π
π΄π .
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )