复杂场景下多模态情感预测:数学建模与算法设计
2026 年中国研究生数学建模竞赛 E 题
复杂场景下多模态情感预测的数学建模与算法设计
1. 项目概述
本项目基于 CMU-MOSEI 英文多模态情感数据,研究复杂场景下文本(Text)、语音(Audio)和视觉(Vision)三种模态的联合情感预测。
赛题主要包含三个逐层递进的问题:
- Q1:多模态情感特征提取与时序对齐
- Q2:模态局部缺失条件下的鲁棒情感预测
- Q3:可解释性多模态情感预测
整体逻辑可以概括为:
[ \boxed{ \text{原始视频} \rightarrow \text{多模态特征} \rightarrow \text{跨模态时序对齐} \rightarrow \text{缺失感知鲁棒融合} \rightarrow \text{情感预测} \rightarrow \text{预测解释} } ]
本项目现阶段不预先指定某一种方案为最终方案,而采用:
[ \boxed{\text{多方案全部实现 + 统一实验协议 + 系统对比}} ]
最终根据:
- 基础预测性能;
- 缺失情况下的鲁棒性;
- 解释可信度;
- 模型复杂度;
- 可复现性;
共同确定最终建模方案。
2. 赛题核心任务
2.1 Q1:多模态特征提取与时序对齐
从原始视频中分别提取:
- 文本语义特征;
- 语音情感特征;
- 视觉情感特征;
并解决三种模态:
- 采样频率不同;
- 序列长度不同;
- 时间尺度不同;
造成的时序不一致问题。
Q1 不以最终情感分类为主要目标,而应产生:
[ \boxed{ \text{结构化、可追溯、可核验的多模态时序特征} } ]
同时记录:
- 原始视频 ID;
- 模态;
- 时间位置;
- 有效长度;
- padding;
- 原始视频时间映射;
- 特征维度;
- 对齐关系。
2.2 Q2:局部模态缺失下的鲁棒情感预测
题目中的“模态缺失”不是简单的:
整个 Audio / Vision / Text 完全不存在。
而是:
[ \boxed{\text{一个或多个模态中出现随机的连续局部缺失区间}} ]
例如:
[ X^A = [a_1,a_2,a_3, \underbrace{0,0,0,0}_{\text{局部缺失}}, a_8,\ldots]. ]
模型需要同时预测:
情感极性
[ \hat c \in { \text{Negative}, \text{Neutral}, \text{Positive} } ]
连续情感强度
[ \hat y\in[-3,3] ]
同时研究:
[ \boxed{ \text{缺失模态类型} + \text{缺失位置} + \text{缺失长度} } ]
对预测性能的影响。
2.3 Q3:可解释性情感预测
在三模态信息完整条件下,不仅预测:
[ (\hat c,\hat y) ]
还需要解释:
- 主要参考了哪种模态;
- 不同模态的作用程度;
- 哪些局部文本、音频和视觉片段最重要;
- 关键证据对应原始视频的什么时间位置。
因此 Q3 的核心不是简单画 Attention,而是:
[ \boxed{ \text{定位证据} + \text{量化贡献} + \text{验证解释} + \text{回溯原始素材} } ]
3. 数据集说明
3.1 附件 1:100 条原始视频
附件 1 包含从 CMU-MOSEI 中筛选的:
[ 100 ]
条英文视频样本。
共有:
[ 37 ]
个 video_id 子文件夹。
样本由:
video_id + clip_id
共同唯一确定。
视频长度范围:
[ 2.648s \sim 34.567s ]
配套文件:
label-100.xlsx
包含字段:
| 字段 | 含义 |
|---|---|
| video_id | 原视频编号 |
| clip_id | 视频片段编号 |
| text | 英文转写文本 |
| label | 连续情感强度 |
| annotation | Negative / Neutral / Positive |
连续情感标签:
[ y\in[-3,3] ]
定义:
[ [-3,0)\Rightarrow Negative ]
[ 0\Rightarrow Neutral ]
[ (0,3]\Rightarrow Positive ]
注意:
[ \boxed{0\text{ 只属于 Neutral}} ]
4. 附件 2:标准化多模态特征
附件 2 包含两套特征:
aligned_50.pkl
unaligned_50.pkl
每套包含约:
[ 4850 ]
条有效样本,并按照:
train
valid
test
划分。
4.1 Pickle 数据组织方式
整体结构为:
data
├── train
│ ├── id
│ ├── raw_text
│ ├── text
│ ├── text_bert
│ ├── audio
│ ├── vision
│ ├── annotations
│ ├── classification_labels
│ ├── regression_labels
│ └── ...
├── valid
└── test
正确读取方式:
data["train"]["audio"][j]
而不是:
data["train"][j]["audio"]
5. aligned / unaligned 数据格式
5.1 aligned_50.pkl
[ X^T\in\mathbb R^{N\times50\times768} ]
[ X^A\in\mathbb R^{N\times50\times74} ]
[ X^V\in\mathbb R^{N\times50\times35} ]
三个模态已经被组织为:
[ 50 ]
个相互对应的位置。
可以理解为:
[ T_i\leftrightarrow A_i\leftrightarrow V_i. ]
5.2 unaligned_50.pkl
文本:
[ X^T\in\mathbb R^{N\times50\times768} ]
音频:
[ X^A\in\mathbb R^{N\times500\times74} ]
视觉:
[ X^V\in\mathbb R^{N\times500\times35} ]
同时提供:
audio_lengths
vision_lengths
记录实际有效长度。
因此 unaligned 数据保留了:
[ \boxed{ \text{文本较粗语义轴} + \text{音频细粒度时序} + \text{视觉细粒度时序} } ]
这为学习式 Cross-Attention 对齐提供了空间。
6. 附件 3:模态局部缺失测试集
附件 3:
- 无标签;
- 已处理多模态特征;
- 包含随机局部模态缺失;
- 缺失区间表现为连续位置全部置零。
用于 Q2 最终预测。
7. 附件 4:可解释性专项测试集
附件 4:
- 无标签;
- 三模态完整;
- 同时提供原始真实场景视频;
- 特征组织形式与附件 2 一致。
用于:
[ \boxed{ \text{预测} + \text{模态贡献} + \text{关键证据定位} } ]
可以根据 Q1 保存的时间映射将关键位置重新定位到:
- 原始文本;
- 音频时间段;
- 视频关键帧。
8. 统一数学符号
一个样本记为:
[ \mathcal X_i= (X_i^T,X_i^A,X_i^V) ]
其中:
- (T):Text
- (A):Audio
- (V):Vision
原始模态序列:
[ X^m= [x_1^m,\ldots,x_{L_m}^m], \qquad m\in{T,A,V}. ]
8.1 基础符号
| 符号 | 含义 |
|---|---|
| (x_j^m) | 模态 (m) 原始第 (j) 个位置特征 |
| (\tilde x_i^m) | 软对齐到文本位置 (i) 后的特征 |
| (h_i^m) | 模态编码器生成的隐藏表示 |
| (M_j^m) | 原始位置是否有效 |
| (A_{ij}^m) | Cross-Attention 对齐权重 |
| (R_i^m) | 对齐后的局部模态可靠度 |
| (\alpha_i^m) | 融合时模态权重 |
| (z_i) | 第 (i) 个位置的多模态融合表示 |
| (\beta_i) | 第 (i) 个位置对整段情绪的时间重要性 |
| (g) | 整条视频最终表示 |
| (\hat y) | 情感强度预测 |
| (\hat c) | 情感极性预测 |
Q1
@Q1.md
可能使用的开源库、工具与模型
| 类别 | 工具 / 模型 | 用途 | 优先级 |
|---|---|---|---|
| 深度学习 | PyTorch | 网络、训练、Cross-Attention | 必须 |
| GPU | CUDA / cuDNN | GPU 加速 | 必须 |
| 数值计算 | NumPy | 数值运算 | 必须 |
| 数据 | Pandas | CSV / XLSX / 表格 | 必须 |
| 科学计算 | SciPy | Pearson / 统计分析 | 必须 |
| ML | scikit-learn | F1、Accuracy、MAE 等 | 必须 |
| 配置 | PyYAML | 实验配置 | 推荐 |
| 配置 | Hydra / OmegaConf | 多实验管理 | 推荐 |
| 日志 | TensorBoard | loss / metric | 推荐 |
| 日志 | MLflow | 系统实验管理 | 可选 |
| 调参 | Optuna | 超参数搜索 | 推荐 |
| 视频 | FFmpeg | 拆音频、转码 | 必须 |
| 视频 | OpenCV | 抽帧与时间映射 | 必须 |
| 文本 | Hugging Face Transformers | BERT 等 | 高 |
| 文本 | BERT-base-uncased | 768 维表示 | 高 |
| ASR/时间戳 | WhisperX | 词级时间定位 | 推荐 |
| 强制对齐 | Montreal Forced Aligner | transcript-audio alignment | 推荐 |
| 音频 | openSMILE | 声学/韵律特征 | 高 |
| 音频 | librosa | 基础信号处理 | 推荐 |
| 音频 | torchaudio | PyTorch 音频 | 推荐 |
| 音频 PLM | wav2vec 2.0 | 深层语音表示 | 可选 |
| 音频 PLM | WavLM | 深层语音表示 | 可选 |
| 音频 PLM | HuBERT | 深层语音表示 | 可选 |
| 视觉 | OpenFace 2.0 | AU / gaze / pose | 高 |
| 视觉 | MediaPipe | landmark / blendshape | 可选 |
| 视觉 | torchvision | 图像处理 | 推荐 |
| 多模态 | CMU-MultimodalSDK | MOSEI 数据参考 | 高 |
| 多模态 | MultiBench | baseline / multimodal benchmark | 高 |
| 解释 | Captum | IG / Occlusion / Ablation | 高 |
| 解释 | SHAP | Shapley 解释 | 可选 |
| 可视化 | Matplotlib | 曲线 / 热图 | 必须 |
| 可视化 | Plotly | 交互可视化 | 可选 |
| 统计 | statsmodels | 回归 / 显著性分析 | 推荐 |
模型候选汇总
| 模块 | 候选 |
|---|---|
| Text Encoder | BERT |
| Audio Feature | openSMILE |
| Audio Encoder | BiGRU / Transformer |
| Audio PLM | Wav2Vec2 / WavLM |
| Vision Feature | OpenFace |
| Vision Encoder | BiGRU / Transformer |
| Alignment | Hard Alignment |
| Alignment | Fixed Window |
| Alignment | Cross-Attention |
| Fusion | Concatenation |
| Fusion | TFN |
| Fusion | LMF |
| Fusion | MAG |
| Fusion | MulT |
| Representation | MISA |
| Missing Model | MMIN |
| Missing Model | CMAD |
| Missing Model | P-RMF |
| Experts | EMOE |
| Diffusion | HyperEF |
| Factorization | FUSE-Net |
| Distribution Alignment | CaReFlow |
| Robust Representation | CmIR |
| Attribution | Integrated Gradients |
| Attribution | Occlusion |
| Attribution | Feature Ablation |
| Attribution | SHAP |
参考文献
下表优先使用论文官方会议/期刊入口。
README 中记录 DOI、ACL Anthology ID、arXiv ID 或官方论文标题,避免引用二手博客作为正式论文来源。
| # | 文献 | 会议/期刊 | 年份 | 与本项目关系 | 官方标识 |
|---|---|---|---|---|---|
| 1 | Zhang et al., Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects | Expert Systems with Applications | 2024 | 多模态情感综述 | DOI: 10.1016/j.eswa.2023.121692 |
| 2 | Bagher Zadeh et al., Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph | ACL | 2018 | CMU-MOSEI | ACL: P18-1208 |
| 3 | Vaswani et al., Attention Is All You Need | NeurIPS | 2017 | Transformer / Attention | NeurIPS 2017 |
| 4 | Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding | NAACL | 2019 | 文本特征 | ACL: N19-1423 |
| 5 | Eyben et al., openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor | ACM Multimedia | 2010 | 音频特征 | ACM MM 2010 |
| 6 | Baltrušaitis et al., OpenFace 2.0: Facial Behavior Analysis Toolkit | FG | 2018 | 视觉特征 | DOI: 10.1109/FG.2018.00019 |
| 7 | Bain et al., WhisperX: Time-Accurate Speech Transcription of Long-Form Audio | arXiv | 2023 | 词级时间戳 | arXiv:2303.00747 |
| 8 | Tsai et al., Multimodal Transformer for Unaligned Multimodal Language Sequences | ACL | 2019 | unaligned + Cross-Attention | ACL: P19-1656 |
| 9 | Rahman et al., Integrating Multimodal Information in Large Pretrained Transformers | ACL | 2020 | MAG-BERT / 文本中心融合 | DOI:10.18653/v1/2020.acl-main.214 |
| 10 | Zadeh et al., Tensor Fusion Network for Multimodal Sentiment Analysis | EMNLP | 2017 | TFN baseline | ACL: D17-1115 |
| 11 | Liu et al., Efficient Low-rank Multimodal Fusion With Modality-Specific Factors | ACL | 2018 | LMF baseline | ACL: P18-1209 |
| 12 | Hazarika et al., MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis | ACM MM | 2020 | shared / private representation | DOI:10.1145/3394171.3413678 |
| 13 | Neverova et al., ModDrop: Adaptive Multi-Modal Gesture Recognition | TPAMI | 2016 | 模态随机丢失 | DOI:10.1109/TPAMI.2015.2461544 |
| 14 | Pham et al., Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities | AAAI | 2019 | 跨模态重构 | DOI:10.1609/aaai.v33i01.33016892 |
| 15 | Zhao et al., Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities | ACL-IJCNLP | 2021 | MMIN | ACL:2021.acl-long.203 |
| 16 | 王楠、王淇、欧阳丹彤,《基于知识蒸馏与动态调整机制的多模态情感分析模型》 | 计算机学报 | 2025 | AUMDF / 动态权重 / 缺失 | 48(8):1923–1942 |
| 17 | Fang et al., EMOE: Modality-Specific Enhanced Dynamic Emotion Experts | CVPR | 2025 | Mixture of Experts | CVPR 2025 |
| 18 | Zhuang et al., CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities | ICCV | 2025 | 缺失模态蒸馏 | ICCV 2025 |
| 19 | Zhu et al., Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data | ACL | 2025 | P-RMF / 不确定性 | DOI:10.18653/v1/2025.acl-long.1075 |
| 20 | Qiu et al., Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition | CVPR | 2026 | Diffusion / uncertainty | CVPR 2026 |
| 21 | Yang & Li, Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis | CVPR | 2026 | shared/specific/noise + reconstruction | CVPR 2026 |
| 22 | Mai & Han, Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective | ACL | 2026 | CmIR / causal invariance | DOI:10.18653/v1/2026.acl-long.2119 |
| 23 | Mai & Han, CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion | CVPR | 2026 | 模态分布对齐 | arXiv:2602.19140 |
| 24 | Wan et al., Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation | ACL | 2026 | 关键证据定位 | DOI:10.18653/v1/2026.acl-long.2012 |
| 25 | Sundararajan et al., Axiomatic Attribution for Deep Networks | ICML | 2017 | Integrated Gradients | PMLR 70 |
| 26 | Jain & Wallace, Attention is not Explanation | NAACL | 2019 | Attention 解释局限 | DOI:10.18653/v1/N19-1357 |
| 27 | Lundberg & Lee, A Unified Approach to Interpreting Model Predictions | NeurIPS | 2017 | SHAP | NeurIPS 2017 |
| 28 | Adebayo et al., Sanity Checks for Saliency Maps | NeurIPS | 2018 | 解释可信性验证 | NeurIPS 2018 |
| 29 | Zeiler & Fergus, Visualizing and Understanding Convolutional Networks | ECCV | 2014 | Occlusion 思想 | DOI:10.1007/978-3-319-10590-1_53 |
| 30 | Wachter et al., Counterfactual Explanations without Opening the Black Box | arXiv / HILDA | 2017 | 反事实解释 | arXiv:1711.00399 |
最终原则
本项目当前不采用:
“先选一个看起来高级的模型,然后证明它最好。”
而采用:
[ \boxed{ \text{提出多个机制假设} \rightarrow \text{设计公平实验} \rightarrow \text{观察数据} \rightarrow \text{解释规律} \rightarrow \text{确定最终模型} } ]
重点不只是得到最高指标,还要回答:
- 为什么某种对齐方式更好?
- 哪种模态最怕局部缺失?
- 缺失的位置是否比缺失比例更重要?
- Cross-Attention 能否帮助判断缺失信息的重要程度?
- 动态可靠度是否真的提高鲁棒性?
- 重构和“降低信任”哪种策略更有效?
- Attention 权重是否真的对应模型决策依据?
- 哪种解释方法最能经受反事实验证?
最终目标是形成:
[ \boxed{ \text{有性能} + \text{有机制} + \text{有数学分析} + \text{有可解释性} } ]
的完整多模态情感预测建模方案。