A multimodal model that predicts whether a conversation is negative, neutral, or positive using text, audio, and video.
The project uses pretrained feature extractors for each modality and a cross-modal Transformer to combine them. It includes training, evaluation, experiment tracking, checkpoint recovery, automated tests, and a Gradio demo.
Results
The final model was evaluated on the complete held-out MELD test split.
Metric Result Test samples 2,610 Accuracy 64.87% Macro-F1 0.6222 Negative F1 0.6172 Neutral F1 0.7183 Positive F1 0.5312
Neutral sentiment performs best. Positive sentiment is the most difficult class and is the main area for future improvement.
Architecture
The model follows four simple steps:
Extract features
RoBERTa represents the spoken words.
Wav2Vec2 represents the voice and speech signal.
CLIP represents the video frames.
Process each modality
A small Transformer learns patterns over time separately for text, audio, and video.
Share information between modalities
Text, audio, and video are compared with one another.
For example, the text may sound positive while the voice or facial expression suggests otherwise.
This step produces one improved representation for each modality.
Predict sentiment
The three representations are combined.
A classification layer predicts negative, neutral, or positive.
Training approach
The final training setup uses:
weighted cross-entropy for class imbalance;
label smoothing and dropout to reduce overfitting;
modality dropout to handle missing or weak signals;
only the final RoBERTa encoder layer unfrozen;
separate learning rates for the classification head and the remaining model;
gradient clipping, cosine learning-rate decay, and early stopping;
model selection based on validation macro-F1.
MLflow records training metrics and artifacts. A complete checkpoint is saved after every epoch, allowing interrupted Kaggle or Colab runs to resume.
Project structure
Installation
Create and activate a virtual environment, then install the training dependencies:
For the Gradio application and raw video inference:
FFmpeg must also be installed and available on PATH.