Scikit-Learn Full Crash Course - Python Machine Learning

NeuralNineAbout 7 min readAug 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Scikit-learn (sklearn): A Python toolkit for traditional machine learning (regression, classification, clustering, PCA) without deep learning.
  • Supervised Learning: Training a model on labeled data (input features and corresponding output labels).
  • Unsupervised Learning: Discovering patterns in unlabeled data (e.g., clustering).
  • Classification: Predicting a discrete class label (e.g., malignant or benign).
  • Regression: Predicting a continuous numerical value (e.g., house price).
  • Pre-processing: Transforming raw data into a suitable format for machine learning models (e.g., scaling, encoding).
  • Feature Scaling: Transforming numerical features to a similar range (e.g., using StandardScaler or MinMaxScaler).
  • Feature Encoding: Converting categorical features into numerical representations (e.g., using OrdinalEncoder or OneHotEncoder).
  • Train-Test Split: Dividing the dataset into training and testing sets to evaluate model generalization.
  • Cross-Validation: A technique for evaluating model performance by training and testing on multiple subsets of the data.
  • Hyperparameter Tuning: Optimizing model performance by searching for the best combination of hyperparameter values.
  • Pipeline: A sequence of data transformations and a final estimator (classifier or regressor) that can be applied in a streamlined manner.

Setting Up the Environment

  • Installing required packages: scikit-learn, numpy, pandas, matplotlib.
  • Using a package manager like pip or uv.
  • Creating a virtual environment using venv or virtualenv.
  • Using Jupyter Lab for interactive Python notebooks to run individual code cells.

Basic Scikit-learn Workflow Example

  • Loading the breast cancer dataset using sklearn.datasets.load_breast_cancer.
  • Splitting the data into training and testing sets using sklearn.model_selection.train_test_split.
  • Scaling the data using sklearn.preprocessing.StandardScaler.
  • Training a K-Nearest Neighbors classifier using sklearn.neighbors.KNeighborsClassifier.
  • Evaluating the model's performance using the score method.
  • Making predictions on new instances using the predict method.

Working with Datasets

  • Using sklearn.datasets to load, fetch, or generate datasets.
  • load_* functions load datasets from the package itself (e.g., load_breast_cancer, load_iris).
  • fetch_* functions download datasets from the internet (e.g., fetch_california_housing, fetch_openml).
  • make_* functions generate synthetic datasets (e.g., make_blobs, make_moons).
  • Parameters like return_X_y=True to get features (X) and target (y) separately.
  • Parameter as_frame=True to load the dataset as a Pandas DataFrame.
  • Using fetch_openml to access datasets like "Titanic" or "car".
  • Using make_blobs to generate clustered data points.
  • Using make_moons to generate data with a crescent shape.
  • Using random_state for reproducible data generation.

Splitting the Data

  • Using sklearn.model_selection.train_test_split to split data into training and testing sets.
  • Specifying test_size to control the proportion of data used for testing (e.g., test_size=0.2 for 20% testing).
  • Understanding the importance of splitting data to avoid overfitting.
  • Using numpy and matplotlib to visualize the distribution of classes in the training and testing sets.
  • Using sklearn.model_selection.StratifiedShuffleSplit to ensure a similar class distribution in both sets.
  • Example:
    sss = StratifiedShuffleSplit(n_splits=1, test_size=0.2)
    for train_index, test_index in sss.split(X, y):
        X_train, X_test = X[train_index], X[test_index]
        y_train, y_test = y[train_index], y[test_index]
    

Pre-processing the Data

  • Using sklearn.preprocessing for scaling, transforming, and encoding data.
  • Scaling:
    • StandardScaler: Scales data to have zero mean and unit variance.
      • Formula: (x - mean) / standard_deviation
    • MinMaxScaler: Scales data to a range between 0 and 1.
      • Formula: (x - min) / (max - min)
    • Using fit_transform on the training data and transform on the testing data to avoid data leakage.
  • Encoding:
    • OrdinalEncoder: Encodes categorical features with ordinal relationships (e.g., "small," "medium," "large") into numerical values.
      • Specifying the order of categories using the categories parameter.
    • OneHotEncoder: Encodes categorical features without ordinal relationships (e.g., "red," "green," "blue") into binary columns.
      • Using handle_unknown='ignore' to handle unknown categories during transformation.
      • Using sparse_output=False to get a dense array instead of a sparse matrix.
      • Using get_feature_names_out to get the names of the new columns.
    • Example:
      encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
      encoded_values = encoder.fit_transform(data[['occupation', 'race']])
      new_columns = encoder.get_feature_names_out(['occupation', 'race'])
      df_encoded = pd.DataFrame(encoded_values, columns=new_columns, index=data.index)
      data_final = pd.concat([data.drop(['occupation', 'race'], axis=1), df_encoded], axis=1)
      

Classification

  • Importing classifiers from sklearn.neighbors, sklearn.linear_model, sklearn.tree, sklearn.svm, sklearn.ensemble, and sklearn.naive_bayes.
  • Creating an instance of the classifier (e.g., KNeighborsClassifier()).
  • Training the classifier using fit(X_train_scaled, y_train).
  • Evaluating the classifier using score(X_test_scaled, y_test).
  • Making predictions using predict(X_test_scaled).
  • Using predict_proba to get probability estimates for each class (if supported by the classifier).
  • Examples:
    • KNeighborsClassifier
    • LogisticRegression
    • DecisionTreeClassifier
    • SVC (Support Vector Classifier)
    • RandomForestClassifier
    • GaussianNB (Gaussian Naive Bayes)
  • Understanding the importance of scaling data for distance-based classifiers (e.g., K-Nearest Neighbors, Logistic Regression, Support Vector Machines).
  • Understanding that tree-based models (e.g., Decision Trees, Random Forests) and Naive Bayes are not scale-sensitive.

Regression

  • Importing regressors from sklearn.linear_model, sklearn.neighbors, sklearn.svm, and sklearn.ensemble.
  • Creating an instance of the regressor (e.g., LinearRegression()).
  • Training the regressor using fit(X_train_scaled, y_train).
  • Evaluating the regressor using score(X_test_scaled, y_test) (returns R-squared).
  • Making predictions using predict(X_test_scaled).
  • Examples:
    • LinearRegression
    • Lasso (L1 regularization)
    • Ridge (L2 regularization)
    • ElasticNet (combination of L1 and L2 regularization)
    • KNeighborsRegressor
    • SVR (Support Vector Regressor)
    • RandomForestRegressor
    • DecisionTreeRegressor

Unsupervised Learning: Clustering

  • Using sklearn.cluster for clustering algorithms.
  • Generating data using make_blobs or make_moons.
  • Scaling the data using StandardScaler.
  • K-Means Clustering:
    • Importing KMeans from sklearn.cluster.
    • Specifying the number of clusters using n_clusters.
    • Training the model using fit(X_scaled).
    • Getting cluster labels using labels_.
  • DBScan Clustering:
    • Importing DBSCAN from sklearn.cluster.
    • Specifying the epsilon value for the distance using eps.
    • Training the model using fit(X_scaled).
    • Getting cluster labels using labels_.
  • Understanding that clustering is about recognizing distinct groups rather than finding the "truth."

Unsupervised Learning: Dimensionality Reduction (PCA)

  • Using sklearn.decomposition.PCA for principal component analysis.
  • Reducing the number of features while preserving variance.
  • Specifying the number of components using n_components.
  • Training the PCA model using fit(X_train).
  • Transforming the data using transform(X_train) and transform(X_test).
  • Using explained_variance_ratio_ to see how much variance each component explains.
  • Example:
    pca = PCA(n_components=10)
    X_train_reduced = pca.fit_transform(X_train)
    X_test_reduced = pca.transform(X_test)
    print(pca.explained_variance_ratio_)
    

Model Evaluation Metrics

  • Using sklearn.metrics for evaluating model performance.
  • Classification Metrics:
    • accuracy_score: Proportion of correctly classified instances.
    • precision_score: Proportion of true positives among predicted positives.
    • recall_score: Proportion of true positives among actual positives.
    • f1_score: Harmonic mean of precision and recall.
  • Regression Metrics:
    • r2_score: R-squared (coefficient of determination), measures the proportion of variance explained by the model.
    • mean_absolute_error: Average absolute difference between predicted and actual values.
    • mean_squared_error: Average squared difference between predicted and actual values.
    • root_mean_squared_error: Square root of the mean squared error.
  • Example:
    from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
    y_pred = clf.predict(X_test_scaled)
    print(f"Accuracy: {accuracy_score(y_test, y_pred)}")
    print(f"Precision: {precision_score(y_test, y_pred)}")
    print(f"Recall: {recall_score(y_test, y_pred)}")
    print(f"F1 Score: {f1_score(y_test, y_pred)}")
    

Cross-Validation

  • Using sklearn.model_selection.cross_val_score to evaluate model performance using cross-validation.
  • Splitting the data into cv folds.
  • Training and testing the model on different combinations of folds.
  • Getting multiple scores and taking the average.
  • Specifying the scoring metric using the scoring parameter (e.g., scoring='precision').
  • Example:
    from sklearn.model_selection import cross_val_score
    scores = cross_val_score(clf, X_scaled, y, cv=5, scoring='accuracy')
    print(f"Cross-validation scores: {scores}")
    print(f"Average cross-validation score: {np.mean(scores)}")
    

Hyperparameter Tuning

  • Using sklearn.model_selection.GridSearchCV to find the best combination of hyperparameters.
  • Defining a parameter grid with the hyperparameters to tune and their possible values.
  • Creating an instance of GridSearchCV with the classifier, parameter grid, and cross-validation folds.
  • Training the grid search model using fit(X_train, y_train).
  • Getting the best parameters using best_params_.
  • Getting the best estimator using best_estimator_.
  • Evaluating the best estimator on the testing set.
  • Example:
    from sklearn.model_selection import GridSearchCV
    param_grid = {
        'n_estimators': [50, 100, 200],
        'min_samples_split': [2, 5],
        'max_depth': [None, 5, 10]
    }
    grid = GridSearchCV(RandomForestClassifier(n_jobs=-1), param_grid, cv=3)
    grid.fit(X_train, y_train)
    print(f"Best parameters: {grid.best_params_}")
    best_clf = grid.best_estimator_
    print(f"Test score: {best_clf.score(X_test, y_test)}")
    

Pipelines

  • Using sklearn.pipeline.Pipeline to chain multiple data transformations and a final estimator.
  • Creating a pipeline with a list of steps, where each step is a tuple containing a name and a transformer or estimator.
  • Training the pipeline using fit(X_train, y_train).
  • Evaluating the pipeline using score(X_test, y_test).
  • Example:
    from sklearn.pipeline import Pipeline
    pipe = Pipeline([
        ('scaler', StandardScaler()),
        ('pca', PCA(n_components=10)),
        ('forest', RandomForestClassifier())
    ])
    pipe.fit(X_train, y_train)
    print(f"Test score: {pipe.score(X_test, y_test)}")
    

Conclusion

Scikit-learn provides a comprehensive set of tools for traditional machine learning tasks. This crash course covered the essential steps of a machine learning workflow, including data loading, splitting, pre-processing, model training, evaluation, hyperparameter tuning, and pipeline creation. By understanding these concepts and utilizing the various modules and functions within scikit-learn, you can effectively build and deploy machine learning models for a wide range of applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.