题目描述:https://www.kaggle.com/c/santander-customer-satisfaction
简单总结:一堆匿名属性;label是0/1;目标是最大化AUC。
第一次尝试:
特征:
由于时间比较充裕,直接用了暴力搜索提取较好的特征:
#!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! better use a RFC or GBC as the clf
#!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! because the final predict model are those two
#!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! we should select better feature for RFC or GBC, not for LR
clf = LogisticRegression(class_weight='balanced', penalty='l2', n_jobs=-1)
selectedFeaInds=GreedyFeatureAdd(clf, trainX, trainY, scoreType="auc", goodFeatures=[], maxFeaNum=150)
joblib.dump(selectedFeaInds, 'modelPersistence/selectedFeaInds.pkl')
#selectedFeaInds=joblib.load('modelPersistence/selectedFeaInds.pkl')
trainX=trainX[:,selectedFeaInds]
testX=testX[:,selectedFeaInds