监督学习笔记 - Datacamp

最新推荐文章于 2021-12-17 16:57:03 发布

AliangYeah

最新推荐文章于 2021-12-17 16:57:03 发布

阅读量269

点赞数

分类专栏： Python 文章标签： python 机器学习

本文链接：https://blog.youkuaiyun.com/AliangYeah/article/details/120300870

版权

这篇博客探讨了监督学习中的分类与回归问题，包括k-Nearest Neighbors、Logistic Regression、简单线性回归、套索回归和岭回归等模型。同时，介绍了模型性能评估方法，如训练集与测试集分离、交叉验证以及ROC曲线绘制。此外，还讨论了性能优化技术，如GridSearchCV和RandomizedSearchCV，以及模型泛化、过拟合和欠拟合的概念。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

分类还是回归？

一、分类

1. k-Nearest Neighbors k-近邻

# Import KNeighborsClassifier from sklearn.neighbors
from sklearn.neighbors import KNeighborsClassifier

# Create arrays for the features and the response variable
y = df['party'].values
X = df.drop('party', axis=1).values

# Create a k-NN classifier with 6 neighbors
knn = KNeighborsClassifier(n_neighbors = 6)

# Fit the classifier to the data
knn.fit(X,y)

# Predict the labels for the training data X: y_pred
y_pred = knn.predict(X)

# Predict and print the label for the new data point X_new
new_prediction = knn.predict(X_new)
print("Prediction: {}".format(new_prediction))

2. Logistic regression 逻辑斯蒂回归

二、回归

1. Basic Linear Regression 简单线性回归

# Import LinearRegression
from sklearn.linear_model import LinearRegression

# Create the regressor: reg
reg = LinearRegression()

# Create the prediction space
prediction_space = np.linspace(min(X_fertility), max(X_fertility)).reshape(-1,1)

# Fit the model to the data
reg.fit(X_fertility, y)

# Compute predictions over the prediction space: y_pred
y_pred = reg.predict(prediction_space)

# Print R^2 
print(reg.score(X_fertility, y))

# Plot regression line
plt.plot(prediction_space, y_pred, color='black', linewidth=3)
plt.show()

2. Regularization I: Lasso Regression 套索回归

# Import Lasso
from sklearn.linear_model import Lasso

# Instantiate a lasso regressor: lasso
lasso = Lasso(alpha=0.4, normalize=True)

# Fit the regressor to the data
l

最低0.47元/天解锁文章