from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
# Load the iris dataset
iris = load_iris()
# Split the data into training and test sets
X_train, X_test, y_train, y_test = train_test_split(iris.data, iris.target, test_size=0.2)
# Create a KNeighborsClassifier object
knn = KNeighborsClassifier(n_neighbors=5)
# Fit the model to the training data
knn.fit(X_train, y_train)
# Make predictions on the test data
y_pred = knn.predict(X_test)
# Evaluate the model's accuracy
print("Accuracy:", knn.score(X_test, y_test))scikit-learn is Python's general-purpose machine learning library. It
covers classification, regression, clustering and preprocessing, and its
models share one interface: create a model, call fit() with training
data, then predict() on new rows. It also ships small datasets,
cross-validation and scoring functions. People use it to learn machine
learning and to build baseline models to compare with
XGBoost or LightGBM. This page
is an online scikit-learn compiler: the code runs in your browser, so you
can try it without installing anything.
Run the example first, then paste any snippet below into a new cell to try it.
load_iris() loads the iris dataset: 150 flowers of three species.
iris.data is a 150 by 4 array of sepal and petal sizes in centimeters,
and iris.target holds the species as 0, 1 or 2. train_test_split()
shuffles the rows and sets 20%, 30 flowers, aside for testing.
KNeighborsClassifier(n_neighbors=5) labels a flower with the most
common species among its 5 closest training flowers.
fit() stores the training data, predict() labels the test flowers,
and score() returns the share it gets right.
The split is random, so the accuracy changes between runs, usually from
0.9 to 1.0. Pass random_state=0 to train_test_split() to get the
same split every time.
k-nearest neighbors compares distances, so a column with large numbers
drowns out the others. In the wine dataset, proline runs from 278 to
1680 and hue from 0.48 to 1.71. StandardScaler rescales every column
to mean 0 and standard deviation 1, and make_pipeline() chains it with
the model:
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0, stratify=y)
raw = KNeighborsClassifier().fit(X_train, y_train)
scaled = make_pipeline(StandardScaler(), KNeighborsClassifier()).fit(X_train, y_train)
print("without scaling:", round(raw.score(X_test, y_test), 3))
print("with scaling: ", round(scaled.score(X_test, y_test), 3))
Scaling raises the test accuracy from 0.667 to 0.933.
Accuracy hides which classes get mixed up. confusion_matrix() counts
each pair of true and predicted class, and classification_report()
gives precision and recall per class:
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
data = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(
data.data, data.target, random_state=0, stratify=data.target)
model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=data.target_names))
Matrix rows are true classes and columns are predictions, in the order malignant, benign: 3 of the 53 malignant tumors are predicted benign.
Regression models predict a number with the same calls. The diabetes dataset's target measures disease progression one year after 10 baseline measurements:
from sklearn.datasets import load_diabetes
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split
X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)
for model in [LinearRegression(), RandomForestRegressor(n_estimators=100, random_state=0)]:
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print(type(model).__name__,
"R2:", round(r2_score(y_test, y_pred), 3),
"MAE:", round(mean_absolute_error(y_test, y_pred), 1))
On this split the linear model wins, with an R² of 0.359 against 0.219 for the random forest.
GridSearchCV scores every combination of settings with 5-fold
cross-validation on the training data and refits the best one. A
pipeline setting is named <step>__<parameter>, where the step is the
class name in lower case:
from sklearn.datasets import load_wine
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0, stratify=y)
pipe = make_pipeline(StandardScaler(), KNeighborsClassifier())
grid = GridSearchCV(pipe, {"kneighborsclassifier__n_neighbors": [1, 3, 5, 7, 9, 11],
"kneighborsclassifier__weights": ["uniform", "distance"]}, cv=5)
grid.fit(X_train, y_train)
print(grid.best_params_)
print("cross-validated accuracy:", round(grid.best_score_, 3))
print("test accuracy:", round(grid.score(X_test, y_test), 3))
LogisticRegression() on the breast cancer data hits
its 100-iteration limit and warns ConvergenceWarning: lbfgs failed to converge. The scaled pipeline above does not.GridSearchCV refit the scaler on
each training fold, so the held-out fold never affects the scaling.load_* functions read data files that ship inside scikit-learn.
fetch_* functions, such as fetch_california_housing(), download
their data first.