Online Scikit-Learn Compiler

Run scikit-learn code in your browser. Train classifiers and regressors, scale data in pipelines and tune models.

Python
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier

# Load the iris dataset
iris = load_iris()

# Split the data into training and test sets
X_train, X_test, y_train, y_test = train_test_split(iris.data, iris.target, test_size=0.2)

# Create a KNeighborsClassifier object
knn = KNeighborsClassifier(n_neighbors=5)

# Fit the model to the training data
knn.fit(X_train, y_train)

# Make predictions on the test data
y_pred = knn.predict(X_test)

# Evaluate the model's accuracy
print("Accuracy:", knn.score(X_test, y_test))
Accuracy: 0.9666666666666667

scikit-learn is Python's general-purpose machine learning library. It covers classification, regression, clustering and preprocessing, and its models share one interface: create a model, call fit() with training data, then predict() on new rows. It also ships small datasets, cross-validation and scoring functions. People use it to learn machine learning and to build baseline models to compare with XGBoost or LightGBM. This page is an online scikit-learn compiler: the code runs in your browser, so you can try it without installing anything. Run the example first, then paste any snippet below into a new cell to try it.

What the example does

load_iris() loads the iris dataset: 150 flowers of three species. iris.data is a 150 by 4 array of sepal and petal sizes in centimeters, and iris.target holds the species as 0, 1 or 2. train_test_split() shuffles the rows and sets 20%, 30 flowers, aside for testing. KNeighborsClassifier(n_neighbors=5) labels a flower with the most common species among its 5 closest training flowers. fit() stores the training data, predict() labels the test flowers, and score() returns the share it gets right.

The split is random, so the accuracy changes between runs, usually from 0.9 to 1.0. Pass random_state=0 to train_test_split() to get the same split every time.

Scale features in a pipeline

k-nearest neighbors compares distances, so a column with large numbers drowns out the others. In the wine dataset, proline runs from 278 to 1680 and hue from 0.48 to 1.71. StandardScaler rescales every column to mean 0 and standard deviation 1, and make_pipeline() chains it with the model:

from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0, stratify=y)

raw = KNeighborsClassifier().fit(X_train, y_train)
scaled = make_pipeline(StandardScaler(), KNeighborsClassifier()).fit(X_train, y_train)
print("without scaling:", round(raw.score(X_test, y_test), 3))
print("with scaling:   ", round(scaled.score(X_test, y_test), 3))

Scaling raises the test accuracy from 0.667 to 0.933.

See which classes the model confuses

Accuracy hides which classes get mixed up. confusion_matrix() counts each pair of true and predicted class, and classification_report() gives precision and recall per class:

from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

data = load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(
    data.data, data.target, random_state=0, stratify=data.target)

model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_train, y_train)
y_pred = model.predict(X_test)

print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=data.target_names))

Matrix rows are true classes and columns are predictions, in the order malignant, benign: 3 of the 53 malignant tumors are predicted benign.

Fit a regression model

Regression models predict a number with the same calls. The diabetes dataset's target measures disease progression one year after 10 baseline measurements:

from sklearn.datasets import load_diabetes
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
from sklearn.model_selection import train_test_split

X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

for model in [LinearRegression(), RandomForestRegressor(n_estimators=100, random_state=0)]:
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)
    print(type(model).__name__,
          "R2:", round(r2_score(y_test, y_pred), 3),
          "MAE:", round(mean_absolute_error(y_test, y_pred), 1))

On this split the linear model wins, with an R² of 0.359 against 0.219 for the random forest.

GridSearchCV scores every combination of settings with 5-fold cross-validation on the training data and refits the best one. A pipeline setting is named <step>__<parameter>, where the step is the class name in lower case:

from sklearn.datasets import load_wine
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0, stratify=y)

pipe = make_pipeline(StandardScaler(), KNeighborsClassifier())
grid = GridSearchCV(pipe, {"kneighborsclassifier__n_neighbors": [1, 3, 5, 7, 9, 11],
                           "kneighborsclassifier__weights": ["uniform", "distance"]}, cv=5)
grid.fit(X_train, y_train)

print(grid.best_params_)
print("cross-validated accuracy:", round(grid.best_score_, 3))
print("test accuracy:", round(grid.score(X_test, y_test), 3))

Good to know

  • Without scaling, LogisticRegression() on the breast cancer data hits its 100-iteration limit and warns ConvergenceWarning: lbfgs failed to converge. The scaled pipeline above does not.
  • In a pipeline, cross-validation and GridSearchCV refit the scaler on each training fold, so the held-out fold never affects the scaling.
  • load_* functions read data files that ship inside scikit-learn. fetch_* functions, such as fetch_california_housing(), download their data first.