下面使用一个简单的电影评论数据集来进行情感分析,目标是判断评论是积极的还是消极的。整个流程包括数据加载、文本预处理、特征提取、模型训练和评估。

import numpy as np
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report
import gensim
from gensim.models import Word2Vec
import nltk
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
import string

# 下载必要的nltk数据
nltk.download('punkt')
nltk.download('stopwords')

# 文本预处理函数
def preprocess_text(text):
    # 转换为小写
    text = text.lower()
    # 去除标点符号
    text = text.translate(str.maketrans('', '', string.punctuation))
    # 分词
    tokens = word_tokenize(text)
    # 去除停用词
    stop_words = set(stopwords.words('english'))
    tokens = [token for token in tokens if token not in stop_words]
    return tokens

# 加载示例数据
data = {
    'text': [
        'This movie is amazing! I really enjoyed it.',
        'The plot was so boring. I couldn\'t even finish it.',
        'Great acting and beautiful cinematography.',
        'Waste of time. Terrible movie.'
    ],
    'label': [1, 0, 1, 0]
}
df = pd.DataFrame(data)

# 文本预处理
df['tokens'] = df['text'].apply(preprocess_text)

# 特征提取 - TF-IDF
tfidf_vectorizer = TfidfVectorizer(tokenizer=lambda x: x, lowercase=False)
tfidf_features = tfidf_vectorizer.fit_transform(df['tokens'])

# 特征提取 - 词向量
model = Word2Vec(df['tokens'], min_count=1)
def get_vector(tokens):
    vectors = [model.wv[token] for token in tokens if token in model.wv]
    if not vectors:
        return np.zeros(model.vector_size)
    return np.mean(vectors, axis=0)

w2v_features = np.array([get_vector(tokens) for tokens in df['tokens']])

# 划分训练集和测试集(以TF-IDF特征为例)
X_train_tfidf, X_test_tfidf, y_train, y_test = train_test_split(tfidf_features, df['label'], test_size=0.2, random_state=42)

# 模型训练 - 逻辑回归
model_tfidf = LogisticRegression()
model_tfidf.fit(X_train_tfidf, y_train)

# 模型评估 - TF-IDF
y_pred_tfidf = model_tfidf.predict(X_test_tfidf)
accuracy_tfidf = accuracy_score(y_test, y_pred_tfidf)
print("TF-IDF Accuracy:", accuracy_tfidf)
print("TF-IDF Classification Report:")
print(classification_report(y_test, y_pred_tfidf))

# 划分训练集和测试集(以词向量特征为例)
X_train_w2v, X_test_w2v, y_train, y_test = train_test_split(w2v_features, df['label'], test_size=0.2, random_state=42)

# 模型训练 - 逻辑回归
model_w2v = LogisticRegression()
model_w2v.fit(X_train_w2v, y_train)

# 模型评估 - 词向量
y_pred_w2v = model_w2v.predict(X_test_w2v)
accuracy_w2v = accuracy_score(y_test, y_pred_w2v)
print("Word2Vec Accuracy:", accuracy_w2v)
print("Word2Vec Classification Report:")
print(classification_report(y_test, y_pred_w2v))

代码解释

文本预处理:

preprocess_text 函数将文本转换为小写,去除标点符号,分词并去除停用词。

使用 apply 方法将预处理函数应用到数据集中的每个文本。

特征提取:

TF - IDF:使用 TfidfVectorizer 从预处理后的文本中提取 TF - IDF 特征。

词向量:使用 Word2Vec 训练词向量模型,然后为每个文本计算平均词向量。

模型训练和评估:

使用 train_test_split 函数将数据集划分为训练集和测试集。

使用逻辑回归模型进行训练,并使用 accuracy_score 和 classification_report 评估模型性能。

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐