【SVM class_weight 】MAP - Charting Student Math Misunderstandings

技术与健康

于 2025-07-19 17:48:36 发布

阅读量30

点赞数

CC 4.0 BY-SA版权

文章标签：支持向量机机器学习人工智能

本文为博主原创文章，未经博主允许不得转载。

本文链接：https://blog.youkuaiyun.com/Practicer2015/article/details/149467737

针对数据不平衡问题，用调整类别权重的方式来处理数据不平衡问题，同时使用支持向量机（SVM）模型进行训练。

我们通过使用 class_weight=‘balanced’ 来调整模型对少数类别的关注度。

import pandas as pd
import string
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC
from sklearn.metrics import classification_report, confusion_matrix, ConfusionMatrixDisplay
import matplotlib.pyplot as plt

# Step 1: Load the dataset
file_path = '/kaggle/input/map-charting-student-math-misunderstandings/train.csv'  # 修改为实际文件路径
data = pd.read_csv(file_path)

# Step 2: Clean the student explanation text (remove punctuation and lower case)
def clean_text(text):
    text = text.lower()  # Convert to lower case
    text = ''.join([char for char in text if char not in string.punctuation])  # Remove punctuation
    return text

# Apply the cleaning function to the 'StudentExplanation' column
data['cleaned_explanation'] = data['StudentExplanation'].apply(clean_text)

# Step 3: Feature extraction using TF-IDF
vectorizer = TfidfVectorizer(stop_words='english', max_features=5000)
X = vectorizer.fit_transform(data['cleaned_explanation'])

# Step 4: Prepare labels (Misconception column)
# We will predict if the explanation contains a misconception or not
data['Misconception'] = data['Misconception'].fillna('No_Misconception')

# Convert labels to binary: 'No_Misconception' -> 0, any other label -> 1
y = data['Misconception'].apply(lambda x: 0 if x == 'No_Misconception' else 1)

# Step 5: Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Step 6: Train an SVM model with class_weight='balanced' to handle the data imbalance
svm_model_weighted = SVC(kernel='linear', class_weight='balanced', random_state=42)
svm_model_weighted.fit(X_train, y_train)

# Step 7: Make predictions
y_pred_svm_weighted = svm_model_weighted.predict(X_test)

# Step 8: Evaluate the model
print(classification_report(y_test, y_pred_svm_weighted))

# Step 9: Plot confusion matrix
cm_weighted = confusion_matrix(y_test, y_pred_svm_weighted)

# Use ConfusionMatrixDisplay to display the confusion matrix
disp = ConfusionMatrixDisplay(confusion_matrix=cm_weighted, display_labels=['No Misconception', 'Misconception'])
disp.plot(cmap=plt.cm.Blues)
plt.title('SVM Model with Balanced Class Weight Confusion Matrix')
plt.show()

 precision    recall  f1-score   support

           0       0.91      0.75      0.82      5277
           1       0.56      0.82      0.67      2063

    accuracy                           0.77      7340
   macro avg       0.74      0.78      0.74      7340
weighted avg       0.81      0.77      0.78      7340

请添加图片描述