Files
modeling_zhaocui/deep_learning/README.md
T

652 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 复杂场景下多模态情感预测:数学建模与算法设计
> 2026 年中国研究生数学建模竞赛 E 题
> **复杂场景下多模态情感预测的数学建模与算法设计**
---
## 1. 项目概述
本项目基于 CMU-MOSEI 英文多模态情感数据,研究复杂场景下文本(Text)、语音(Audio)和视觉(Vision)三种模态的联合情感预测。
赛题主要包含三个逐层递进的问题:
1. **Q1:多模态情感特征提取与时序对齐**
2. **Q2:模态局部缺失条件下的鲁棒情感预测**
3. **Q3:可解释性多模态情感预测**
整体逻辑可以概括为:
\[
\boxed{
\text{原始视频}
\rightarrow
\text{多模态特征}
\rightarrow
\text{跨模态时序对齐}
\rightarrow
\text{缺失感知鲁棒融合}
\rightarrow
\text{情感预测}
\rightarrow
\text{预测解释}
}
\]
本项目现阶段不预先指定某一种方案为最终方案,而采用:
\[
\boxed{\text{多方案全部实现 + 统一实验协议 + 系统对比}}
\]
最终根据:
- 基础预测性能;
- 缺失情况下的鲁棒性;
- 解释可信度;
- 模型复杂度;
- 可复现性;
共同确定最终建模方案。
---
# 2. 赛题核心任务
## 2.1 Q1:多模态特征提取与时序对齐
从原始视频中分别提取:
- 文本语义特征;
- 语音情感特征;
- 视觉情感特征;
并解决三种模态:
- 采样频率不同;
- 序列长度不同;
- 时间尺度不同;
造成的时序不一致问题。
Q1 不以最终情感分类为主要目标,而应产生:
\[
\boxed{
\text{结构化、可追溯、可核验的多模态时序特征}
}
\]
同时记录:
- 原始视频 ID;
- 模态;
- 时间位置;
- 有效长度;
- padding;
- 原始视频时间映射;
- 特征维度;
- 对齐关系。
---
## 2.2 Q2:局部模态缺失下的鲁棒情感预测
题目中的“模态缺失”不是简单的:
> 整个 Audio / Vision / Text 完全不存在。
而是:
\[
\boxed{\text{一个或多个模态中出现随机的连续局部缺失区间}}
\]
例如:
\[
X^A =
[a_1,a_2,a_3,
\underbrace{0,0,0,0}_{\text{局部缺失}},
a_8,\ldots].
\]
模型需要同时预测:
### 情感极性
\[
\hat c \in
\{
\text{Negative},
\text{Neutral},
\text{Positive}
\}
\]
### 连续情感强度
\[
\hat y\in[-3,3]
\]
同时研究:
\[
\boxed{
\text{缺失模态类型}
+
\text{缺失位置}
+
\text{缺失长度}
}
\]
对预测性能的影响。
---
## 2.3 Q3:可解释性情感预测
在三模态信息完整条件下,不仅预测:
\[
(\hat c,\hat y)
\]
还需要解释:
1. 主要参考了哪种模态;
2. 不同模态的作用程度;
3. 哪些局部文本、音频和视觉片段最重要;
4. 关键证据对应原始视频的什么时间位置。
因此 Q3 的核心不是简单画 Attention,而是:
\[
\boxed{
\text{定位证据}
+
\text{量化贡献}
+
\text{验证解释}
+
\text{回溯原始素材}
}
\]
---
# 3. 数据集说明
---
## 3.1 附件 1:100 条原始视频
附件 1 包含从 CMU-MOSEI 中筛选的:
\[
100
\]
条英文视频样本。
共有:
\[
37
\]
个 `video_id` 子文件夹。
样本由:
```text
video_id + clip_id
```
共同唯一确定。
视频长度范围:
\[
2.648s \sim 34.567s
\]
配套文件:
```text
label-100.xlsx
```
包含字段:
| 字段 | 含义 |
|---|---|
| video_id | 原视频编号 |
| clip_id | 视频片段编号 |
| text | 英文转写文本 |
| label | 连续情感强度 |
| annotation | Negative / Neutral / Positive |
连续情感标签:
\[
y\in[-3,3]
\]
定义:
\[
[-3,0)\Rightarrow Negative
\]
\[
0\Rightarrow Neutral
\]
\[
(0,3]\Rightarrow Positive
\]
注意:
\[
\boxed{0\text{ 只属于 Neutral}}
\]
---
# 4. 附件 2:标准化多模态特征
附件 2 包含两套特征:
```text
aligned_50.pkl
unaligned_50.pkl
```
每套包含约:
\[
4850
\]
条有效样本,并按照:
```python
train
valid
test
```
划分。
---
## 4.1 Pickle 数据组织方式
整体结构为:
```text
data
├── train
│ ├── id
│ ├── raw_text
│ ├── text
│ ├── text_bert
│ ├── audio
│ ├── vision
│ ├── annotations
│ ├── classification_labels
│ ├── regression_labels
│ └── ...
├── valid
└── test
```
正确读取方式:
```python
data["train"]["audio"][j]
```
而不是:
```python
data["train"][j]["audio"]
```
---
# 5. aligned / unaligned 数据格式
## 5.1 aligned_50.pkl
\[
X^T\in\mathbb R^{N\times50\times768}
\]
\[
X^A\in\mathbb R^{N\times50\times74}
\]
\[
X^V\in\mathbb R^{N\times50\times35}
\]
三个模态已经被组织为:
\[
50
\]
个相互对应的位置。
可以理解为:
\[
T_i\leftrightarrow A_i\leftrightarrow V_i.
\]
---
## 5.2 unaligned_50.pkl
文本:
\[
X^T\in\mathbb R^{N\times50\times768}
\]
音频:
\[
X^A\in\mathbb R^{N\times500\times74}
\]
视觉:
\[
X^V\in\mathbb R^{N\times500\times35}
\]
同时提供:
```text
audio_lengths
vision_lengths
```
记录实际有效长度。
因此 unaligned 数据保留了:
\[
\boxed{
\text{文本较粗语义轴}
+
\text{音频细粒度时序}
+
\text{视觉细粒度时序}
}
\]
这为学习式 Cross-Attention 对齐提供了空间。
---
# 6. 附件 3:模态局部缺失测试集
附件 3:
- 无标签;
- 已处理多模态特征;
- 包含随机局部模态缺失;
- 缺失区间表现为连续位置全部置零。
用于 Q2 最终预测。
---
# 7. 附件 4:可解释性专项测试集
附件 4:
- 无标签;
- 三模态完整;
- 同时提供原始真实场景视频;
- 特征组织形式与附件 2 一致。
用于:
\[
\boxed{
\text{预测}
+
\text{模态贡献}
+
\text{关键证据定位}
}
\]
可以根据 Q1 保存的时间映射将关键位置重新定位到:
- 原始文本;
- 音频时间段;
- 视频关键帧。
---
# 8. 统一数学符号
一个样本记为:
\[
\mathcal X_i=
(X_i^T,X_i^A,X_i^V)
\]
其中:
- \(T\):Text
- \(A\):Audio
- \(V\):Vision
原始模态序列:
\[
X^m=
[x_1^m,\ldots,x_{L_m}^m],
\qquad
m\in\{T,A,V\}.
\]
---
## 8.1 基础符号
| 符号 | 含义 |
|---|---|
| \(x_j^m\) | 模态 \(m\) 原始第 \(j\) 个位置特征 |
| \(\tilde x_i^m\) | 软对齐到文本位置 \(i\) 后的特征 |
| \(h_i^m\) | 模态编码器生成的隐藏表示 |
| \(M_j^m\) | 原始位置是否有效 |
| \(A_{ij}^m\) | Cross-Attention 对齐权重 |
| \(R_i^m\) | 对齐后的局部模态可靠度 |
| \(\alpha_i^m\) | 融合时模态权重 |
| \(z_i\) | 第 \(i\) 个位置的多模态融合表示 |
| \(\beta_i\) | 第 \(i\) 个位置对整段情绪的时间重要性 |
| \(g\) | 整条视频最终表示 |
| \(\hat y\) | 情感强度预测 |
| \(\hat c\) | 情感极性预测 |
---
## Q1
@Q1.md
---
# 可能使用的开源库、工具与模型
| 类别 | 工具 / 模型 | 用途 | 优先级 |
|---|---|---|---|
| 深度学习 | PyTorch | 网络、训练、Cross-Attention | 必须 |
| GPU | CUDA / cuDNN | GPU 加速 | 必须 |
| 数值计算 | NumPy | 数值运算 | 必须 |
| 数据 | Pandas | CSV / XLSX / 表格 | 必须 |
| 科学计算 | SciPy | Pearson / 统计分析 | 必须 |
| ML | scikit-learn | F1、Accuracy、MAE 等 | 必须 |
| 配置 | PyYAML | 实验配置 | 推荐 |
| 配置 | Hydra / OmegaConf | 多实验管理 | 推荐 |
| 日志 | TensorBoard | loss / metric | 推荐 |
| 日志 | MLflow | 系统实验管理 | 可选 |
| 调参 | Optuna | 超参数搜索 | 推荐 |
| 视频 | FFmpeg | 拆音频、转码 | 必须 |
| 视频 | OpenCV | 抽帧与时间映射 | 必须 |
| 文本 | Hugging Face Transformers | BERT 等 | 高 |
| 文本 | BERT-base-uncased | 768 维表示 | 高 |
| ASR/时间戳 | WhisperX | 词级时间定位 | 推荐 |
| 强制对齐 | Montreal Forced Aligner | transcript-audio alignment | 推荐 |
| 音频 | openSMILE | 声学/韵律特征 | 高 |
| 音频 | librosa | 基础信号处理 | 推荐 |
| 音频 | torchaudio | PyTorch 音频 | 推荐 |
| 音频 PLM | wav2vec 2.0 | 深层语音表示 | 可选 |
| 音频 PLM | WavLM | 深层语音表示 | 可选 |
| 音频 PLM | HuBERT | 深层语音表示 | 可选 |
| 视觉 | OpenFace 2.0 | AU / gaze / pose | 高 |
| 视觉 | MediaPipe | landmark / blendshape | 可选 |
| 视觉 | torchvision | 图像处理 | 推荐 |
| 多模态 | CMU-MultimodalSDK | MOSEI 数据参考 | 高 |
| 多模态 | MultiBench | baseline / multimodal benchmark | 高 |
| 解释 | Captum | IG / Occlusion / Ablation | 高 |
| 解释 | SHAP | Shapley 解释 | 可选 |
| 可视化 | Matplotlib | 曲线 / 热图 | 必须 |
| 可视化 | Plotly | 交互可视化 | 可选 |
| 统计 | statsmodels | 回归 / 显著性分析 | 推荐 |
---
# 模型候选汇总
| 模块 | 候选 |
|---|---|
| Text Encoder | BERT |
| Audio Feature | openSMILE |
| Audio Encoder | BiGRU / Transformer |
| Audio PLM | Wav2Vec2 / WavLM |
| Vision Feature | OpenFace |
| Vision Encoder | BiGRU / Transformer |
| Alignment | Hard Alignment |
| Alignment | Fixed Window |
| Alignment | Cross-Attention |
| Fusion | Concatenation |
| Fusion | TFN |
| Fusion | LMF |
| Fusion | MAG |
| Fusion | MulT |
| Representation | MISA |
| Missing Model | MMIN |
| Missing Model | CMAD |
| Missing Model | P-RMF |
| Experts | EMOE |
| Diffusion | HyperEF |
| Factorization | FUSE-Net |
| Distribution Alignment | CaReFlow |
| Robust Representation | CmIR |
| Attribution | Integrated Gradients |
| Attribution | Occlusion |
| Attribution | Feature Ablation |
| Attribution | SHAP |
---
# 参考文献
> 下表优先使用论文官方会议/期刊入口。
> README 中记录 DOI、ACL Anthology ID、arXiv ID 或官方论文标题,避免引用二手博客作为正式论文来源。
| # | 文献 | 会议/期刊 | 年份 | 与本项目关系 | 官方标识 |
|---|---|---|---:|---|---|
| 1 | Zhang et al., *Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects* | Expert Systems with Applications | 2024 | 多模态情感综述 | DOI: 10.1016/j.eswa.2023.121692 |
| 2 | Bagher Zadeh et al., *Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph* | ACL | 2018 | CMU-MOSEI | ACL: P18-1208 |
| 3 | Vaswani et al., *Attention Is All You Need* | NeurIPS | 2017 | Transformer / Attention | NeurIPS 2017 |
| 4 | Devlin et al., *BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding* | NAACL | 2019 | 文本特征 | ACL: N19-1423 |
| 5 | Eyben et al., *openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor* | ACM Multimedia | 2010 | 音频特征 | ACM MM 2010 |
| 6 | Baltrušaitis et al., *OpenFace 2.0: Facial Behavior Analysis Toolkit* | FG | 2018 | 视觉特征 | DOI: 10.1109/FG.2018.00019 |
| 7 | Bain et al., *WhisperX: Time-Accurate Speech Transcription of Long-Form Audio* | arXiv | 2023 | 词级时间戳 | arXiv:2303.00747 |
| 8 | Tsai et al., *Multimodal Transformer for Unaligned Multimodal Language Sequences* | ACL | 2019 | unaligned + Cross-Attention | ACL: P19-1656 |
| 9 | Rahman et al., *Integrating Multimodal Information in Large Pretrained Transformers* | ACL | 2020 | MAG-BERT / 文本中心融合 | DOI:10.18653/v1/2020.acl-main.214 |
| 10 | Zadeh et al., *Tensor Fusion Network for Multimodal Sentiment Analysis* | EMNLP | 2017 | TFN baseline | ACL: D17-1115 |
| 11 | Liu et al., *Efficient Low-rank Multimodal Fusion With Modality-Specific Factors* | ACL | 2018 | LMF baseline | ACL: P18-1209 |
| 12 | Hazarika et al., *MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis* | ACM MM | 2020 | shared / private representation | DOI:10.1145/3394171.3413678 |
| 13 | Neverova et al., *ModDrop: Adaptive Multi-Modal Gesture Recognition* | TPAMI | 2016 | 模态随机丢失 | DOI:10.1109/TPAMI.2015.2461544 |
| 14 | Pham et al., *Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities* | AAAI | 2019 | 跨模态重构 | DOI:10.1609/aaai.v33i01.33016892 |
| 15 | Zhao et al., *Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities* | ACL-IJCNLP | 2021 | MMIN | ACL:2021.acl-long.203 |
| 16 | 王楠、王淇、欧阳丹彤,《基于知识蒸馏与动态调整机制的多模态情感分析模型》 | 计算机学报 | 2025 | AUMDF / 动态权重 / 缺失 | 48(8):1923–1942 |
| 17 | Fang et al., *EMOE: Modality-Specific Enhanced Dynamic Emotion Experts* | CVPR | 2025 | Mixture of Experts | CVPR 2025 |
| 18 | Zhuang et al., *CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities* | ICCV | 2025 | 缺失模态蒸馏 | ICCV 2025 |
| 19 | Zhu et al., *Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data* | ACL | 2025 | P-RMF / 不确定性 | DOI:10.18653/v1/2025.acl-long.1075 |
| 20 | Qiu et al., *Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition* | CVPR | 2026 | Diffusion / uncertainty | CVPR 2026 |
| 21 | Yang & Li, *Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis* | CVPR | 2026 | shared/specific/noise + reconstruction | CVPR 2026 |
| 22 | Mai & Han, *Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective* | ACL | 2026 | CmIR / causal invariance | DOI:10.18653/v1/2026.acl-long.2119 |
| 23 | Mai & Han, *CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion* | CVPR | 2026 | 模态分布对齐 | arXiv:2602.19140 |
| 24 | Wan et al., *Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation* | ACL | 2026 | 关键证据定位 | DOI:10.18653/v1/2026.acl-long.2012 |
| 25 | Sundararajan et al., *Axiomatic Attribution for Deep Networks* | ICML | 2017 | Integrated Gradients | PMLR 70 |
| 26 | Jain & Wallace, *Attention is not Explanation* | NAACL | 2019 | Attention 解释局限 | DOI:10.18653/v1/N19-1357 |
| 27 | Lundberg & Lee, *A Unified Approach to Interpreting Model Predictions* | NeurIPS | 2017 | SHAP | NeurIPS 2017 |
| 28 | Adebayo et al., *Sanity Checks for Saliency Maps* | NeurIPS | 2018 | 解释可信性验证 | NeurIPS 2018 |
| 29 | Zeiler & Fergus, *Visualizing and Understanding Convolutional Networks* | ECCV | 2014 | Occlusion 思想 | DOI:10.1007/978-3-319-10590-1_53 |
| 30 | Wachter et al., *Counterfactual Explanations without Opening the Black Box* | arXiv / HILDA | 2017 | 反事实解释 | arXiv:1711.00399 |
---
# 最终原则
本项目当前不采用:
> “先选一个看起来高级的模型,然后证明它最好。”
而采用:
\[
\boxed{
\text{提出多个机制假设}
\rightarrow
\text{设计公平实验}
\rightarrow
\text{观察数据}
\rightarrow
\text{解释规律}
\rightarrow
\text{确定最终模型}
}
\]
重点不只是得到最高指标,还要回答:
1. 为什么某种对齐方式更好?
2. 哪种模态最怕局部缺失?
3. 缺失的位置是否比缺失比例更重要?
4. Cross-Attention 能否帮助判断缺失信息的重要程度?
5. 动态可靠度是否真的提高鲁棒性?
6. 重构和“降低信任”哪种策略更有效?
7. Attention 权重是否真的对应模型决策依据?
8. 哪种解释方法最能经受反事实验证?
最终目标是形成:
\[
\boxed{
\text{有性能}
+
\text{有机制}
+
\text{有数学分析}
+
\text{有可解释性}
}
\]
的完整多模态情感预测建模方案。