MACHINE LEARNING • LESSON 11

What Is K-Means?

K-Means is a clustering algorithm that divides data into a chosen number of groups based on how similar the data points are.

THE SIMPLE IDEA

K-Means finds similar data points and puts them into groups.

You tell K-Means how many groups you want by choosing a value for K. K-Means then tries to organize the data into those groups.

01

What Does K-Means Mean?

The name K-Means contains two important ideas: K and Means.

K Number of Clusters

K tells the algorithm how many groups we want to create.

MEANS Center of a Cluster

K-Means uses the average position of the points in a cluster to find its center.

If K = 3, K-Means tries to divide the data into 3 clusters.
02

What Does K Control?

The value of K determines how many groups K-Means should create.

K = 2
● ● ●
● ●
● ● ●
● ●
2 Clusters
K = 3
● ●
● ●
● ●
3 Clusters
K = 4
4 Clusters

So when you write:

KMeans(n_clusters=3)

you are telling the algorithm:

"Create 3 clusters from my data."
03

Simple Customer Example

Imagine an online store has customers with different spending habits.

We record two features:

FEATURE 1 Annual Spending

How much money the customer spends.

FEATURE 2 Number of Purchases

How many purchases the customer makes.

CUSTOMER SPENDING PURCHASES
A ₹1,000 2
B ₹1,200 3
C ₹10,000 15
D ₹11,000 17
E ₹2,000 4

Looking at the numbers, customers A, B, and E are relatively similar.

Customers C and D are also relatively similar.

04

K-Means Can Find These Groups

Suppose we choose:

K = 2

We are asking K-Means to create two groups.

CLUSTER 1 A, B, E

These customers have relatively lower spending and fewer purchases.

CLUSTER 2 C, D

These customers have much higher spending and more purchases.

Customer Data K = 2 2 Clusters
K-Means does not need us to provide "low spender" or "high spender" labels. It discovers the groups from the data.
05

What Is the "Mean"?

The word Means comes from the idea of calculating the average position of the points in a cluster.

For example, imagine three values:

10 + 20 + 30 = 60 ÷ 3 = 20

The mean is 20.

K-Means uses this idea to find the center of each cluster. This center is called a centroid.

CLUSTER
● = Centroid
06

K-Means Is Unsupervised Learning

From the previous page, you learned that unsupervised learning works without predefined labels.

K-Means is an example of unsupervised learning because we don't tell it which customer belongs to which group.

Customer Data K-Means Discovered Clusters
WE PROVIDE Features

Example: spending and number of purchases.

K-MEANS FINDS Groups

Similar customers are placed into clusters.

07

What K-Means Does Not Tell You

There is an important detail beginners often misunderstand.

If K-Means produces:

Cluster 0
Cluster 1

that does not automatically mean:

Cluster 0 = Low-value customers
Cluster 1 = High-value customers

The numbers are simply identifiers.

After clustering, we inspect the data inside each cluster and decide what the groups actually represent.

Cluster 0 is not "better" or "worse" than Cluster 1. The cluster number itself has no business meaning.
08

One More Simple Example

Imagine a school has information about students' study hours and attendance, but no student categories.

STUDENT A 2 hours / 60%
STUDENT B 3 hours / 65%
STUDENT C 7 hours / 90%
STUDENT D 8 hours / 92%

If we choose:

K = 2

K-Means might discover:

CLUSTER 1 A, B

Students with lower study hours and attendance.

CLUSTER 2 C, D

Students with higher study hours and attendance.

REMEMBER THIS

K-Means = Choose K + Find Similar Groups

K-Means is an unsupervised clustering algorithm. You choose how many clusters you want using K, and the algorithm organizes similar data points into those clusters.

K Number of Groups
+
MEANS Cluster Centers
=
K-MEANS Clustering
QUICK CHECK

Check Your Understanding

What is K-Means? An unsupervised clustering algorithm that groups similar data points.
What does K mean? The number of clusters we want the algorithm to create.
What does K = 3 mean? K-Means should create 3 clusters.
What does "Mean" refer to? The average position used to find the center of a cluster.
What is a centroid? The center point representing a cluster.
Is K-Means supervised? No. K-Means is an unsupervised learning algorithm.
NEXT TOPIC

How K-Means Works

Next, we will walk through the actual K-Means process: choosing initial centers, assigning points, moving the centers, and repeating the process.