# 复杂场景下多模态情感预测:数学建模与算法设计 > 2026 年中国研究生数学建模竞赛 E 题 > **复杂场景下多模态情感预测的数学建模与算法设计** --- ## 1. 项目概述 本项目基于 CMU-MOSEI 英文多模态情感数据,研究复杂场景下文本(Text)、语音(Audio)和视觉(Vision)三种模态的联合情感预测。 赛题主要包含三个逐层递进的问题: 1. **Q1:多模态情感特征提取与时序对齐** 2. **Q2:模态局部缺失条件下的鲁棒情感预测** 3. **Q3:可解释性多模态情感预测** 整体逻辑可以概括为: \[ \boxed{ \text{原始视频} \rightarrow \text{多模态特征} \rightarrow \text{跨模态时序对齐} \rightarrow \text{缺失感知鲁棒融合} \rightarrow \text{情感预测} \rightarrow \text{预测解释} } \] 本项目现阶段不预先指定某一种方案为最终方案,而采用: \[ \boxed{\text{多方案全部实现 + 统一实验协议 + 系统对比}} \] 最终根据: - 基础预测性能; - 缺失情况下的鲁棒性; - 解释可信度; - 模型复杂度; - 可复现性; 共同确定最终建模方案。 --- # 2. 赛题核心任务 ## 2.1 Q1:多模态特征提取与时序对齐 从原始视频中分别提取: - 文本语义特征; - 语音情感特征; - 视觉情感特征; 并解决三种模态: - 采样频率不同; - 序列长度不同; - 时间尺度不同; 造成的时序不一致问题。 Q1 不以最终情感分类为主要目标,而应产生: \[ \boxed{ \text{结构化、可追溯、可核验的多模态时序特征} } \] 同时记录: - 原始视频 ID; - 模态; - 时间位置; - 有效长度; - padding; - 原始视频时间映射; - 特征维度; - 对齐关系。 --- ## 2.2 Q2:局部模态缺失下的鲁棒情感预测 题目中的“模态缺失”不是简单的: > 整个 Audio / Vision / Text 完全不存在。 而是: \[ \boxed{\text{一个或多个模态中出现随机的连续局部缺失区间}} \] 例如: \[ X^A = [a_1,a_2,a_3, \underbrace{0,0,0,0}_{\text{局部缺失}}, a_8,\ldots]. \] 模型需要同时预测: ### 情感极性 \[ \hat c \in \{ \text{Negative}, \text{Neutral}, \text{Positive} \} \] ### 连续情感强度 \[ \hat y\in[-3,3] \] 同时研究: \[ \boxed{ \text{缺失模态类型} + \text{缺失位置} + \text{缺失长度} } \] 对预测性能的影响。 --- ## 2.3 Q3:可解释性情感预测 在三模态信息完整条件下,不仅预测: \[ (\hat c,\hat y) \] 还需要解释: 1. 主要参考了哪种模态; 2. 不同模态的作用程度; 3. 哪些局部文本、音频和视觉片段最重要; 4. 关键证据对应原始视频的什么时间位置。 因此 Q3 的核心不是简单画 Attention,而是: \[ \boxed{ \text{定位证据} + \text{量化贡献} + \text{验证解释} + \text{回溯原始素材} } \] --- # 3. 数据集说明 --- ## 3.1 附件 1:100 条原始视频 附件 1 包含从 CMU-MOSEI 中筛选的: \[ 100 \] 条英文视频样本。 共有: \[ 37 \] 个 `video_id` 子文件夹。 样本由: ```text video_id + clip_id ``` 共同唯一确定。 视频长度范围: \[ 2.648s \sim 34.567s \] 配套文件: ```text label-100.xlsx ``` 包含字段: | 字段 | 含义 | |---|---| | video_id | 原视频编号 | | clip_id | 视频片段编号 | | text | 英文转写文本 | | label | 连续情感强度 | | annotation | Negative / Neutral / Positive | 连续情感标签: \[ y\in[-3,3] \] 定义: \[ [-3,0)\Rightarrow Negative \] \[ 0\Rightarrow Neutral \] \[ (0,3]\Rightarrow Positive \] 注意: \[ \boxed{0\text{ 只属于 Neutral}} \] --- # 4. 附件 2:标准化多模态特征 附件 2 包含两套特征: ```text aligned_50.pkl unaligned_50.pkl ``` 每套包含约: \[ 4850 \] 条有效样本,并按照: ```python train valid test ``` 划分。 --- ## 4.1 Pickle 数据组织方式 整体结构为: ```text data ├── train │ ├── id │ ├── raw_text │ ├── text │ ├── text_bert │ ├── audio │ ├── vision │ ├── annotations │ ├── classification_labels │ ├── regression_labels │ └── ... ├── valid └── test ``` 正确读取方式: ```python data["train"]["audio"][j] ``` 而不是: ```python data["train"][j]["audio"] ``` --- # 5. aligned / unaligned 数据格式 ## 5.1 aligned_50.pkl \[ X^T\in\mathbb R^{N\times50\times768} \] \[ X^A\in\mathbb R^{N\times50\times74} \] \[ X^V\in\mathbb R^{N\times50\times35} \] 三个模态已经被组织为: \[ 50 \] 个相互对应的位置。 可以理解为: \[ T_i\leftrightarrow A_i\leftrightarrow V_i. \] --- ## 5.2 unaligned_50.pkl 文本: \[ X^T\in\mathbb R^{N\times50\times768} \] 音频: \[ X^A\in\mathbb R^{N\times500\times74} \] 视觉: \[ X^V\in\mathbb R^{N\times500\times35} \] 同时提供: ```text audio_lengths vision_lengths ``` 记录实际有效长度。 因此 unaligned 数据保留了: \[ \boxed{ \text{文本较粗语义轴} + \text{音频细粒度时序} + \text{视觉细粒度时序} } \] 这为学习式 Cross-Attention 对齐提供了空间。 --- # 6. 附件 3:模态局部缺失测试集 附件 3: - 无标签; - 已处理多模态特征; - 包含随机局部模态缺失; - 缺失区间表现为连续位置全部置零。 用于 Q2 最终预测。 --- # 7. 附件 4:可解释性专项测试集 附件 4: - 无标签; - 三模态完整; - 同时提供原始真实场景视频; - 特征组织形式与附件 2 一致。 用于: \[ \boxed{ \text{预测} + \text{模态贡献} + \text{关键证据定位} } \] 可以根据 Q1 保存的时间映射将关键位置重新定位到: - 原始文本; - 音频时间段; - 视频关键帧。 --- # 8. 统一数学符号 一个样本记为: \[ \mathcal X_i= (X_i^T,X_i^A,X_i^V) \] 其中: - \(T\):Text - \(A\):Audio - \(V\):Vision 原始模态序列: \[ X^m= [x_1^m,\ldots,x_{L_m}^m], \qquad m\in\{T,A,V\}. \] --- ## 8.1 基础符号 | 符号 | 含义 | |---|---| | \(x_j^m\) | 模态 \(m\) 原始第 \(j\) 个位置特征 | | \(\tilde x_i^m\) | 软对齐到文本位置 \(i\) 后的特征 | | \(h_i^m\) | 模态编码器生成的隐藏表示 | | \(M_j^m\) | 原始位置是否有效 | | \(A_{ij}^m\) | Cross-Attention 对齐权重 | | \(R_i^m\) | 对齐后的局部模态可靠度 | | \(\alpha_i^m\) | 融合时模态权重 | | \(z_i\) | 第 \(i\) 个位置的多模态融合表示 | | \(\beta_i\) | 第 \(i\) 个位置对整段情绪的时间重要性 | | \(g\) | 整条视频最终表示 | | \(\hat y\) | 情感强度预测 | | \(\hat c\) | 情感极性预测 | --- ## Q1 @Q1.md --- # 可能使用的开源库、工具与模型 | 类别 | 工具 / 模型 | 用途 | 优先级 | |---|---|---|---| | 深度学习 | PyTorch | 网络、训练、Cross-Attention | 必须 | | GPU | CUDA / cuDNN | GPU 加速 | 必须 | | 数值计算 | NumPy | 数值运算 | 必须 | | 数据 | Pandas | CSV / XLSX / 表格 | 必须 | | 科学计算 | SciPy | Pearson / 统计分析 | 必须 | | ML | scikit-learn | F1、Accuracy、MAE 等 | 必须 | | 配置 | PyYAML | 实验配置 | 推荐 | | 配置 | Hydra / OmegaConf | 多实验管理 | 推荐 | | 日志 | TensorBoard | loss / metric | 推荐 | | 日志 | MLflow | 系统实验管理 | 可选 | | 调参 | Optuna | 超参数搜索 | 推荐 | | 视频 | FFmpeg | 拆音频、转码 | 必须 | | 视频 | OpenCV | 抽帧与时间映射 | 必须 | | 文本 | Hugging Face Transformers | BERT 等 | 高 | | 文本 | BERT-base-uncased | 768 维表示 | 高 | | ASR/时间戳 | WhisperX | 词级时间定位 | 推荐 | | 强制对齐 | Montreal Forced Aligner | transcript-audio alignment | 推荐 | | 音频 | openSMILE | 声学/韵律特征 | 高 | | 音频 | librosa | 基础信号处理 | 推荐 | | 音频 | torchaudio | PyTorch 音频 | 推荐 | | 音频 PLM | wav2vec 2.0 | 深层语音表示 | 可选 | | 音频 PLM | WavLM | 深层语音表示 | 可选 | | 音频 PLM | HuBERT | 深层语音表示 | 可选 | | 视觉 | OpenFace 2.0 | AU / gaze / pose | 高 | | 视觉 | MediaPipe | landmark / blendshape | 可选 | | 视觉 | torchvision | 图像处理 | 推荐 | | 多模态 | CMU-MultimodalSDK | MOSEI 数据参考 | 高 | | 多模态 | MultiBench | baseline / multimodal benchmark | 高 | | 解释 | Captum | IG / Occlusion / Ablation | 高 | | 解释 | SHAP | Shapley 解释 | 可选 | | 可视化 | Matplotlib | 曲线 / 热图 | 必须 | | 可视化 | Plotly | 交互可视化 | 可选 | | 统计 | statsmodels | 回归 / 显著性分析 | 推荐 | --- # 模型候选汇总 | 模块 | 候选 | |---|---| | Text Encoder | BERT | | Audio Feature | openSMILE | | Audio Encoder | BiGRU / Transformer | | Audio PLM | Wav2Vec2 / WavLM | | Vision Feature | OpenFace | | Vision Encoder | BiGRU / Transformer | | Alignment | Hard Alignment | | Alignment | Fixed Window | | Alignment | Cross-Attention | | Fusion | Concatenation | | Fusion | TFN | | Fusion | LMF | | Fusion | MAG | | Fusion | MulT | | Representation | MISA | | Missing Model | MMIN | | Missing Model | CMAD | | Missing Model | P-RMF | | Experts | EMOE | | Diffusion | HyperEF | | Factorization | FUSE-Net | | Distribution Alignment | CaReFlow | | Robust Representation | CmIR | | Attribution | Integrated Gradients | | Attribution | Occlusion | | Attribution | Feature Ablation | | Attribution | SHAP | --- # 参考文献 > 下表优先使用论文官方会议/期刊入口。 > README 中记录 DOI、ACL Anthology ID、arXiv ID 或官方论文标题,避免引用二手博客作为正式论文来源。 | # | 文献 | 会议/期刊 | 年份 | 与本项目关系 | 官方标识 | |---|---|---|---:|---|---| | 1 | Zhang et al., *Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects* | Expert Systems with Applications | 2024 | 多模态情感综述 | DOI: 10.1016/j.eswa.2023.121692 | | 2 | Bagher Zadeh et al., *Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph* | ACL | 2018 | CMU-MOSEI | ACL: P18-1208 | | 3 | Vaswani et al., *Attention Is All You Need* | NeurIPS | 2017 | Transformer / Attention | NeurIPS 2017 | | 4 | Devlin et al., *BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding* | NAACL | 2019 | 文本特征 | ACL: N19-1423 | | 5 | Eyben et al., *openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor* | ACM Multimedia | 2010 | 音频特征 | ACM MM 2010 | | 6 | Baltrušaitis et al., *OpenFace 2.0: Facial Behavior Analysis Toolkit* | FG | 2018 | 视觉特征 | DOI: 10.1109/FG.2018.00019 | | 7 | Bain et al., *WhisperX: Time-Accurate Speech Transcription of Long-Form Audio* | arXiv | 2023 | 词级时间戳 | arXiv:2303.00747 | | 8 | Tsai et al., *Multimodal Transformer for Unaligned Multimodal Language Sequences* | ACL | 2019 | unaligned + Cross-Attention | ACL: P19-1656 | | 9 | Rahman et al., *Integrating Multimodal Information in Large Pretrained Transformers* | ACL | 2020 | MAG-BERT / 文本中心融合 | DOI:10.18653/v1/2020.acl-main.214 | | 10 | Zadeh et al., *Tensor Fusion Network for Multimodal Sentiment Analysis* | EMNLP | 2017 | TFN baseline | ACL: D17-1115 | | 11 | Liu et al., *Efficient Low-rank Multimodal Fusion With Modality-Specific Factors* | ACL | 2018 | LMF baseline | ACL: P18-1209 | | 12 | Hazarika et al., *MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis* | ACM MM | 2020 | shared / private representation | DOI:10.1145/3394171.3413678 | | 13 | Neverova et al., *ModDrop: Adaptive Multi-Modal Gesture Recognition* | TPAMI | 2016 | 模态随机丢失 | DOI:10.1109/TPAMI.2015.2461544 | | 14 | Pham et al., *Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities* | AAAI | 2019 | 跨模态重构 | DOI:10.1609/aaai.v33i01.33016892 | | 15 | Zhao et al., *Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities* | ACL-IJCNLP | 2021 | MMIN | ACL:2021.acl-long.203 | | 16 | 王楠、王淇、欧阳丹彤,《基于知识蒸馏与动态调整机制的多模态情感分析模型》 | 计算机学报 | 2025 | AUMDF / 动态权重 / 缺失 | 48(8):1923–1942 | | 17 | Fang et al., *EMOE: Modality-Specific Enhanced Dynamic Emotion Experts* | CVPR | 2025 | Mixture of Experts | CVPR 2025 | | 18 | Zhuang et al., *CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities* | ICCV | 2025 | 缺失模态蒸馏 | ICCV 2025 | | 19 | Zhu et al., *Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data* | ACL | 2025 | P-RMF / 不确定性 | DOI:10.18653/v1/2025.acl-long.1075 | | 20 | Qiu et al., *Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition* | CVPR | 2026 | Diffusion / uncertainty | CVPR 2026 | | 21 | Yang & Li, *Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis* | CVPR | 2026 | shared/specific/noise + reconstruction | CVPR 2026 | | 22 | Mai & Han, *Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective* | ACL | 2026 | CmIR / causal invariance | DOI:10.18653/v1/2026.acl-long.2119 | | 23 | Mai & Han, *CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion* | CVPR | 2026 | 模态分布对齐 | arXiv:2602.19140 | | 24 | Wan et al., *Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation* | ACL | 2026 | 关键证据定位 | DOI:10.18653/v1/2026.acl-long.2012 | | 25 | Sundararajan et al., *Axiomatic Attribution for Deep Networks* | ICML | 2017 | Integrated Gradients | PMLR 70 | | 26 | Jain & Wallace, *Attention is not Explanation* | NAACL | 2019 | Attention 解释局限 | DOI:10.18653/v1/N19-1357 | | 27 | Lundberg & Lee, *A Unified Approach to Interpreting Model Predictions* | NeurIPS | 2017 | SHAP | NeurIPS 2017 | | 28 | Adebayo et al., *Sanity Checks for Saliency Maps* | NeurIPS | 2018 | 解释可信性验证 | NeurIPS 2018 | | 29 | Zeiler & Fergus, *Visualizing and Understanding Convolutional Networks* | ECCV | 2014 | Occlusion 思想 | DOI:10.1007/978-3-319-10590-1_53 | | 30 | Wachter et al., *Counterfactual Explanations without Opening the Black Box* | arXiv / HILDA | 2017 | 反事实解释 | arXiv:1711.00399 | --- # 最终原则 本项目当前不采用: > “先选一个看起来高级的模型,然后证明它最好。” 而采用: \[ \boxed{ \text{提出多个机制假设} \rightarrow \text{设计公平实验} \rightarrow \text{观察数据} \rightarrow \text{解释规律} \rightarrow \text{确定最终模型} } \] 重点不只是得到最高指标,还要回答: 1. 为什么某种对齐方式更好? 2. 哪种模态最怕局部缺失? 3. 缺失的位置是否比缺失比例更重要? 4. Cross-Attention 能否帮助判断缺失信息的重要程度? 5. 动态可靠度是否真的提高鲁棒性? 6. 重构和“降低信任”哪种策略更有效? 7. Attention 权重是否真的对应模型决策依据? 8. 哪种解释方法最能经受反事实验证? 最终目标是形成: \[ \boxed{ \text{有性能} + \text{有机制} + \text{有数学分析} + \text{有可解释性} } \] 的完整多模态情感预测建模方案。