IMDB Movie Reviews Sentiment Analysis Project by Rabiul IslamIMDB Movie Reviews Sentiment Analysis Project by Rabiul Islam

IMDB Movie Reviews Sentiment Analysis Project

Rabiul Islam

Rabiul Islam

๐ŸŽฌ IMDB Movie Review Sentiment Analysis

A Natural Language Processing and Deep Learning project that performs sentiment classification on 50,000 IMDB movie reviews.
This project compares three different text representation approaches:
TF-IDF + Logistic Regression
Word2Vec + Logistic Regression
BERT-based Sentiment Classification
The goal is to understand how different text representation techniques affect classification accuracy, model performance, computational cost, and error patterns.

๐ŸŽฏ Project Objective

The primary objective is to build and compare multiple sentiment analysis models using the IMDB Movie Reviews dataset.
The project focuses on:
Text preprocessing
Feature extraction
Word embeddings
Transformer-based NLP
Sentiment classification
Model evaluation
Performance comparison
Error analysis

๐Ÿ“Š Dataset

The IMDB Movie Reviews Dataset contains:
50,000 movie reviews
25,000 training reviews
25,000 test reviews
Binary sentiment classification
Positive and negative reviews

Classes

Label Sentiment 0 Negative 1 Positive
The dataset is loaded using the Hugging Face datasets library.

๐Ÿ”„ Project Workflow


๐Ÿงน 1. Data Loading & Preprocessing

The reviews are preprocessed before feature extraction.

Preprocessing Steps

Convert text to lowercase
Remove HTML tags
Remove unnecessary punctuation
Normalize whitespace
Tokenize text where required
For BERT, minimal preprocessing is used because BERT's tokenizer is designed to process natural language text.

Example


Important: BERT uses minimal preprocessing and does not require traditional stopword removal.

๐Ÿ“Œ 2. TF-IDF Sentiment Analysis

TF-IDF

Term Frequency-Inverse Document Frequency represents each review as a numerical vector based on the importance of words within the document collection.

Classifier

Logistic Regression is used as the classification algorithm.

๐Ÿ“Œ 3. Word2Vec Sentiment Analysis

Word2Vec

Word2Vec learns dense vector representations of words based on their contextual relationships.
A Word2Vec model is trained using only the training data to prevent data leakage.

Review Representation

Each review is represented by the average of its word vectors.

The resulting document vectors are used to train a classifier.

๐Ÿ“Œ 4. BERT Sentiment Analysis

BERT

BERT (Bidirectional Encoder Representations from Transformers) is a Transformer-based language model that captures contextual relationships between words.
Unlike TF-IDF and basic Word2Vec representations, BERT considers the context in which words appear.
A pre-trained BERT model is used for sentiment classification.
Example:

The model is fine-tuned on the IMDB training data.

BERT Preprocessing

BERT uses its own tokenizer:

Minimal preprocessing is used before BERT tokenization to preserve important linguistic information.

๐Ÿ“ˆ 5. Model Evaluation

Each model is evaluated on the held-out test dataset.
The following metrics are calculated:
Accuracy
Precision
Recall
F1-Score

Evaluation Function


๐Ÿ“Š 6. Model Comparison

The performance of the three approaches is summarized below.
Model Accuracy Precision Recall F1-Score TF-IDF + Logistic Regression 0.8954 TBD TBD TBD Word2Vec + Logistic Regression 0.8482 TBD TBD TBD BERT 0.9078 TBD TBD 0.9101

Replace the TBD values with the exact values generated by your final notebook.

Current Project Results

Based on the completed experiments:
TF-IDF Accuracy: 89.54%
Word2Vec Accuracy: 84.82%
BERT Accuracy: 90.78%
BERT F1-Score: 91.01%
BERT achieved the strongest overall performance among the tested approaches.

๐Ÿ† Best Performing Model

๐Ÿฅ‡ BERT

BERT achieved the highest classification performance in this project.
Its contextual representation allows the model to understand relationships between words more effectively than traditional TF-IDF and averaged Word2Vec representations.

Performance

Accuracy: 90.78%
F1-Score: 91.01%

๐Ÿ” 7. Model Comparison & Analysis

Key Observations

BERT achieved the best overall performance, demonstrating the advantage of contextual Transformer-based representations.
TF-IDF performed strongly with approximately 89.54% accuracy despite using a relatively simple representation and classifier.
Word2Vec achieved approximately 84.82% accuracy, which was lower than both TF-IDF and BERT.
Averaging Word2Vec embeddings can lose important word-order and contextual information.
TF-IDF is computationally simpler than BERT and can provide strong performance for traditional sentiment classification.
BERT requires significantly more computational resources and training time than TF-IDF or Word2Vec.
BERT can better handle contextual meaning, negation, and relationships between words.
TF-IDF may struggle when sentiment depends heavily on word context rather than individual words or n-grams.
Word2Vec provides dense semantic representations but averaging word vectors can remove important sentence-level information.
For applications where computational resources are limited, TF-IDF can provide an excellent accuracy-to-cost trade-off.

โš ๏ธ Error Analysis

Common sentiment classification errors can occur in reviews containing:
Sarcasm
Mixed positive and negative opinions
Complex sentence structures
Negation
Ambiguous language
Long reviews with multiple opinions
For example, a review may contain many positive words while the overall sentiment is negative because of sarcasm or context.
BERT is generally better suited to these cases because it can model contextual relationships between words.

โฑ๏ธ Computational Comparison

Approach Representation Computational Cost Context Understanding TF-IDF Sparse vectors Low Limited Word2Vec Dense embeddings Medium Moderate BERT Contextual embeddings High Strong

Practical Recommendation

TF-IDF: Best when speed, simplicity, and low computational cost are important.
Word2Vec: Useful for learning semantic word representations and traditional NLP applications.
BERT: Best when maximum classification performance and contextual understanding are priorities.

๐Ÿ”ฌ Reproducibility

Fixed random seeds are used where applicable to improve reproducibility.

For machine learning models:

๐Ÿ› ๏ธ Technologies Used

Python
Pandas
NumPy
Scikit-learn
Gensim
TensorFlow / PyTorch
Hugging Face Transformers
Hugging Face Datasets
Matplotlib
Seaborn
Google Colab

๐Ÿ““ Google Colab

The complete implementation, preprocessing steps, model training, evaluation metrics, comparison table, and final conclusion are available in the Google Colab notebook:
Make sure the sharing permission is:
Anyone with the link โ†’ Viewer

๐Ÿ“ Project Structure


The assignment requires the final submission to be a public Google Colab link. The notebook does not need to be submitted as a separate .ipynb file on the class portal.

๐Ÿ“š Key Learning Outcomes

This project demonstrates practical knowledge of:
Natural Language Processing
Text preprocessing
TF-IDF
Word2Vec
Word embeddings
BERT
Transformer models
Sentiment classification
Logistic Regression
Model evaluation
Accuracy
Precision
Recall
F1-score
Error analysis
Model comparison
Reproducible Machine Learning

๐ŸŽ“ Final Conclusion

This project compared three different approaches for sentiment analysis on the IMDB movie review dataset: TF-IDF, Word2Vec, and BERT. TF-IDF with Logistic Regression achieved strong performance while requiring relatively low computational resources. Word2Vec provided dense semantic representations but performed less effectively when document representations were created by simply averaging word vectors. BERT achieved the highest performance, demonstrating the advantage of contextual Transformer-based representations for understanding complex natural language. The experiments also highlighted an important trade-off between model accuracy and computational cost. While BERT provides stronger contextual understanding and better overall classification performance, TF-IDF remains an attractive solution for applications where simplicity, speed, and resource efficiency are important. Overall, the project demonstrates that the choice of text representation has a significant impact on sentiment classification performance.

๐Ÿ‘จโ€๐Ÿ’ป Author

Rabiul Islam
MSc in CSE โ€” Data Science Machine Learning | Deep Learning | NLP | Artificial Intelligence | Computer Vision
โญ If you find this project useful, consider giving the repository a star.
Like this project

Posted Sep 18, 2026

Developed a sentiment classification model for IMDB reviews using three NLP techniques.