Multimodal Sentiment Analysis Model Development by Dhananjaya PaliwalMultimodal Sentiment Analysis Model Development by Dhananjaya Paliwal

Multimodal Sentiment Analysis Model Development

Dhananjaya Paliwal

Dhananjaya Paliwal

Multimodal Sentiment Analysis

A multimodal model that predicts whether a conversation is negative, neutral, or positive using text, audio, and video.
The project uses pretrained feature extractors for each modality and a cross-modal Transformer to combine them. It includes training, evaluation, experiment tracking, checkpoint recovery, automated tests, and a Gradio demo.

Results

The final model was evaluated on the complete held-out MELD test split.
Metric Result Test samples 2,610 Accuracy 64.87% Macro-F1 0.6222 Negative F1 0.6172 Neutral F1 0.7183 Positive F1 0.5312
Neutral sentiment performs best. Positive sentiment is the most difficult class and is the main area for future improvement.

Architecture

The model follows four simple steps:
Extract features
RoBERTa represents the spoken words.
Wav2Vec2 represents the voice and speech signal.
CLIP represents the video frames.
Process each modality
A small Transformer learns patterns over time separately for text, audio, and video.
Share information between modalities
Text, audio, and video are compared with one another.
For example, the text may sound positive while the voice or facial expression suggests otherwise.
This step produces one improved representation for each modality.
Predict sentiment
The three representations are combined.
A classification layer predicts negative, neutral, or positive.

Training approach

The final training setup uses:
weighted cross-entropy for class imbalance;
label smoothing and dropout to reduce overfitting;
modality dropout to handle missing or weak signals;
only the final RoBERTa encoder layer unfrozen;
separate learning rates for the classification head and the remaining model;
gradient clipping, cosine learning-rate decay, and early stopping;
model selection based on validation macro-F1.
MLflow records training metrics and artifacts. A complete checkpoint is saved after every epoch, allowing interrupted Kaggle or Colab runs to resume.

Project structure


Installation

Create and activate a virtual environment, then install the training dependencies:

For the Gradio application and raw video inference:

FFmpeg must also be installed and available on PATH.

Dataset

This project uses the MELD dataset.
The processed HDF5 file is expected to contain:

The fixed dataset split contains:
Split Samples Train 9,989 Validation 1,108 Test 2,610
Update the local paths in config.json before training.

Training

Run the complete pipeline:

This trains the model, saves the best checkpoint, evaluates the test split, and creates the result plots.
Important outputs:

Running the demo

Set the model and config paths, then start the application.
Windows PowerShell:

Linux or macOS:

Open http://localhost:7860.
The interface supports:
text-only analysis;
uploaded video;
recorded webcam video.
When a video is selected, the text box is disabled and the speech transcript is generated automatically.

Limitations

MELD contains scripted television dialogue, so performance may differ on real meetings or interviews.
Sarcasm and context-dependent dialogue remain difficult.
Positive sentiment has lower recall than the other classes.
Softmax confidence is not calibrated.
Large audio and video sequences make full-dataset evaluation computationally expensive.

Main technologies

PyTorch · Transformers · RoBERTa · Wav2Vec2 · OpenCLIP · MLflow · DVC · HDF5 · Gradio

Acknowledgements

The processed dataset and trained model weights are not stored in this repository because of their size and source licensing.
Like this project

Posted Sep 22, 2026

Developed a multimodal sentiment analysis model using text, audio, and video data.