Key Concepts
- Scikit-learn (sklearn): A Python toolkit for traditional machine learning (regression, classification, clustering, PCA) without deep learning.
- Supervised Learning: Training a model on labeled data (input features and corresponding output labels).
- Unsupervised Learning: Discovering patterns in unlabeled data (e.g., clustering).
- Classification: Predicting a discrete class label (e.g., malignant or benign).
- Regression: Predicting a continuous numerical value (e.g., house price).
- Pre-processing: Transforming raw data into a suitable format for machine learning models (e.g., scaling, encoding).
- Feature Scaling: Transforming numerical features to a similar range (e.g., using StandardScaler or MinMaxScaler).
- Feature Encoding: Converting categorical features into numerical representations (e.g., using OrdinalEncoder or OneHotEncoder).
- Train-Test Split: Dividing the dataset into training and testing sets to evaluate model generalization.
- Cross-Validation: A technique for evaluating model performance by training and testing on multiple subsets of the data.
- Hyperparameter Tuning: Optimizing model performance by searching for the best combination of hyperparameter values.
- Pipeline: A sequence of data transformations and a final estimator (classifier or regressor) that can be applied in a streamlined manner.
Setting Up the Environment
- Installing required packages:
scikit-learn,numpy,pandas,matplotlib. - Using a package manager like
piporuv. - Creating a virtual environment using
venvorvirtualenv. - Using Jupyter Lab for interactive Python notebooks to run individual code cells.
Basic Scikit-learn Workflow Example
- Loading the breast cancer dataset using
sklearn.datasets.load_breast_cancer. - Splitting the data into training and testing sets using
sklearn.model_selection.train_test_split. - Scaling the data using
sklearn.preprocessing.StandardScaler. - Training a K-Nearest Neighbors classifier using
sklearn.neighbors.KNeighborsClassifier. - Evaluating the model's performance using the
scoremethod. - Making predictions on new instances using the
predictmethod.
Working with Datasets
- Using
sklearn.datasetsto load, fetch, or generate datasets. load_*functions load datasets from the package itself (e.g.,load_breast_cancer,load_iris).fetch_*functions download datasets from the internet (e.g.,fetch_california_housing,fetch_openml).make_*functions generate synthetic datasets (e.g.,make_blobs,make_moons).- Parameters like
return_X_y=Trueto get features (X) and target (y) separately. - Parameter
as_frame=Trueto load the dataset as a Pandas DataFrame. - Using
fetch_openmlto access datasets like "Titanic" or "car". - Using
make_blobsto generate clustered data points. - Using
make_moonsto generate data with a crescent shape. - Using
random_statefor reproducible data generation.
Splitting the Data
- Using
sklearn.model_selection.train_test_splitto split data into training and testing sets. - Specifying
test_sizeto control the proportion of data used for testing (e.g.,test_size=0.2for 20% testing). - Understanding the importance of splitting data to avoid overfitting.
- Using
numpyandmatplotlibto visualize the distribution of classes in the training and testing sets. - Using
sklearn.model_selection.StratifiedShuffleSplitto ensure a similar class distribution in both sets. - Example:
sss = StratifiedShuffleSplit(n_splits=1, test_size=0.2) for train_index, test_index in sss.split(X, y): X_train, X_test = X[train_index], X[test_index] y_train, y_test = y[train_index], y[test_index]
Pre-processing the Data
- Using
sklearn.preprocessingfor scaling, transforming, and encoding data. - Scaling:
StandardScaler: Scales data to have zero mean and unit variance.- Formula:
(x - mean) / standard_deviation
- Formula:
MinMaxScaler: Scales data to a range between 0 and 1.- Formula:
(x - min) / (max - min)
- Formula:
- Using
fit_transformon the training data andtransformon the testing data to avoid data leakage.
- Encoding:
OrdinalEncoder: Encodes categorical features with ordinal relationships (e.g., "small," "medium," "large") into numerical values.- Specifying the order of categories using the
categoriesparameter.
- Specifying the order of categories using the
OneHotEncoder: Encodes categorical features without ordinal relationships (e.g., "red," "green," "blue") into binary columns.- Using
handle_unknown='ignore'to handle unknown categories during transformation. - Using
sparse_output=Falseto get a dense array instead of a sparse matrix. - Using
get_feature_names_outto get the names of the new columns.
- Using
- Example:
encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False) encoded_values = encoder.fit_transform(data[['occupation', 'race']]) new_columns = encoder.get_feature_names_out(['occupation', 'race']) df_encoded = pd.DataFrame(encoded_values, columns=new_columns, index=data.index) data_final = pd.concat([data.drop(['occupation', 'race'], axis=1), df_encoded], axis=1)
Classification
- Importing classifiers from
sklearn.neighbors,sklearn.linear_model,sklearn.tree,sklearn.svm,sklearn.ensemble, andsklearn.naive_bayes. - Creating an instance of the classifier (e.g.,
KNeighborsClassifier()). - Training the classifier using
fit(X_train_scaled, y_train). - Evaluating the classifier using
score(X_test_scaled, y_test). - Making predictions using
predict(X_test_scaled). - Using
predict_probato get probability estimates for each class (if supported by the classifier). - Examples:
KNeighborsClassifierLogisticRegressionDecisionTreeClassifierSVC(Support Vector Classifier)RandomForestClassifierGaussianNB(Gaussian Naive Bayes)
- Understanding the importance of scaling data for distance-based classifiers (e.g., K-Nearest Neighbors, Logistic Regression, Support Vector Machines).
- Understanding that tree-based models (e.g., Decision Trees, Random Forests) and Naive Bayes are not scale-sensitive.
Regression
- Importing regressors from
sklearn.linear_model,sklearn.neighbors,sklearn.svm, andsklearn.ensemble. - Creating an instance of the regressor (e.g.,
LinearRegression()). - Training the regressor using
fit(X_train_scaled, y_train). - Evaluating the regressor using
score(X_test_scaled, y_test)(returns R-squared). - Making predictions using
predict(X_test_scaled). - Examples:
LinearRegressionLasso(L1 regularization)Ridge(L2 regularization)ElasticNet(combination of L1 and L2 regularization)KNeighborsRegressorSVR(Support Vector Regressor)RandomForestRegressorDecisionTreeRegressor
Unsupervised Learning: Clustering
- Using
sklearn.clusterfor clustering algorithms. - Generating data using
make_blobsormake_moons. - Scaling the data using
StandardScaler. - K-Means Clustering:
- Importing
KMeansfromsklearn.cluster. - Specifying the number of clusters using
n_clusters. - Training the model using
fit(X_scaled). - Getting cluster labels using
labels_.
- Importing
- DBScan Clustering:
- Importing
DBSCANfromsklearn.cluster. - Specifying the epsilon value for the distance using
eps. - Training the model using
fit(X_scaled). - Getting cluster labels using
labels_.
- Importing
- Understanding that clustering is about recognizing distinct groups rather than finding the "truth."
Unsupervised Learning: Dimensionality Reduction (PCA)
- Using
sklearn.decomposition.PCAfor principal component analysis. - Reducing the number of features while preserving variance.
- Specifying the number of components using
n_components. - Training the PCA model using
fit(X_train). - Transforming the data using
transform(X_train)andtransform(X_test). - Using
explained_variance_ratio_to see how much variance each component explains. - Example:
pca = PCA(n_components=10) X_train_reduced = pca.fit_transform(X_train) X_test_reduced = pca.transform(X_test) print(pca.explained_variance_ratio_)
Model Evaluation Metrics
- Using
sklearn.metricsfor evaluating model performance. - Classification Metrics:
accuracy_score: Proportion of correctly classified instances.precision_score: Proportion of true positives among predicted positives.recall_score: Proportion of true positives among actual positives.f1_score: Harmonic mean of precision and recall.
- Regression Metrics:
r2_score: R-squared (coefficient of determination), measures the proportion of variance explained by the model.mean_absolute_error: Average absolute difference between predicted and actual values.mean_squared_error: Average squared difference between predicted and actual values.root_mean_squared_error: Square root of the mean squared error.
- Example:
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score y_pred = clf.predict(X_test_scaled) print(f"Accuracy: {accuracy_score(y_test, y_pred)}") print(f"Precision: {precision_score(y_test, y_pred)}") print(f"Recall: {recall_score(y_test, y_pred)}") print(f"F1 Score: {f1_score(y_test, y_pred)}")
Cross-Validation
- Using
sklearn.model_selection.cross_val_scoreto evaluate model performance using cross-validation. - Splitting the data into
cvfolds. - Training and testing the model on different combinations of folds.
- Getting multiple scores and taking the average.
- Specifying the scoring metric using the
scoringparameter (e.g.,scoring='precision'). - Example:
from sklearn.model_selection import cross_val_score scores = cross_val_score(clf, X_scaled, y, cv=5, scoring='accuracy') print(f"Cross-validation scores: {scores}") print(f"Average cross-validation score: {np.mean(scores)}")
Hyperparameter Tuning
- Using
sklearn.model_selection.GridSearchCVto find the best combination of hyperparameters. - Defining a parameter grid with the hyperparameters to tune and their possible values.
- Creating an instance of
GridSearchCVwith the classifier, parameter grid, and cross-validation folds. - Training the grid search model using
fit(X_train, y_train). - Getting the best parameters using
best_params_. - Getting the best estimator using
best_estimator_. - Evaluating the best estimator on the testing set.
- Example:
from sklearn.model_selection import GridSearchCV param_grid = { 'n_estimators': [50, 100, 200], 'min_samples_split': [2, 5], 'max_depth': [None, 5, 10] } grid = GridSearchCV(RandomForestClassifier(n_jobs=-1), param_grid, cv=3) grid.fit(X_train, y_train) print(f"Best parameters: {grid.best_params_}") best_clf = grid.best_estimator_ print(f"Test score: {best_clf.score(X_test, y_test)}")
Pipelines
- Using
sklearn.pipeline.Pipelineto chain multiple data transformations and a final estimator. - Creating a pipeline with a list of steps, where each step is a tuple containing a name and a transformer or estimator.
- Training the pipeline using
fit(X_train, y_train). - Evaluating the pipeline using
score(X_test, y_test). - Example:
from sklearn.pipeline import Pipeline pipe = Pipeline([ ('scaler', StandardScaler()), ('pca', PCA(n_components=10)), ('forest', RandomForestClassifier()) ]) pipe.fit(X_train, y_train) print(f"Test score: {pipe.score(X_test, y_test)}")
Conclusion
Scikit-learn provides a comprehensive set of tools for traditional machine learning tasks. This crash course covered the essential steps of a machine learning workflow, including data loading, splitting, pre-processing, model training, evaluation, hyperparameter tuning, and pipeline creation. By understanding these concepts and utilizing the various modules and functions within scikit-learn, you can effectively build and deploy machine learning models for a wide range of applications.
AI summaries can miss context or contain errors. Check important details against the original video.





