建立分批同步基线(基础文件)

This commit is contained in:
2026-09-23 23:24:01 +08:00
commit 7fc76aaafd
70 changed files with 18635 additions and 0 deletions
+652
View File
@@ -0,0 +1,652 @@
# 复杂场景下多模态情感预测:数学建模与算法设计
> 2026 年中国研究生数学建模竞赛 E 题
> **复杂场景下多模态情感预测的数学建模与算法设计**
---
## 1. 项目概述
本项目基于 CMU-MOSEI 英文多模态情感数据,研究复杂场景下文本(Text)、语音(Audio)和视觉(Vision)三种模态的联合情感预测。
赛题主要包含三个逐层递进的问题:
1. **Q1:多模态情感特征提取与时序对齐**
2. **Q2:模态局部缺失条件下的鲁棒情感预测**
3. **Q3:可解释性多模态情感预测**
整体逻辑可以概括为:
\[
\boxed{
\text{原始视频}
\rightarrow
\text{多模态特征}
\rightarrow
\text{跨模态时序对齐}
\rightarrow
\text{缺失感知鲁棒融合}
\rightarrow
\text{情感预测}
\rightarrow
\text{预测解释}
}
\]
本项目现阶段不预先指定某一种方案为最终方案,而采用:
\[
\boxed{\text{多方案全部实现 + 统一实验协议 + 系统对比}}
\]
最终根据:
- 基础预测性能;
- 缺失情况下的鲁棒性;
- 解释可信度;
- 模型复杂度;
- 可复现性;
共同确定最终建模方案。
---
# 2. 赛题核心任务
## 2.1 Q1:多模态特征提取与时序对齐
从原始视频中分别提取:
- 文本语义特征;
- 语音情感特征;
- 视觉情感特征;
并解决三种模态:
- 采样频率不同;
- 序列长度不同;
- 时间尺度不同;
造成的时序不一致问题。
Q1 不以最终情感分类为主要目标,而应产生:
\[
\boxed{
\text{结构化、可追溯、可核验的多模态时序特征}
}
\]
同时记录:
- 原始视频 ID;
- 模态;
- 时间位置;
- 有效长度;
- padding;
- 原始视频时间映射;
- 特征维度;
- 对齐关系。
---
## 2.2 Q2:局部模态缺失下的鲁棒情感预测
题目中的“模态缺失”不是简单的:
> 整个 Audio / Vision / Text 完全不存在。
而是:
\[
\boxed{\text{一个或多个模态中出现随机的连续局部缺失区间}}
\]
例如:
\[
X^A =
[a_1,a_2,a_3,
\underbrace{0,0,0,0}_{\text{局部缺失}},
a_8,\ldots].
\]
模型需要同时预测:
### 情感极性
\[
\hat c \in
\{
\text{Negative},
\text{Neutral},
\text{Positive}
\}
\]
### 连续情感强度
\[
\hat y\in[-3,3]
\]
同时研究:
\[
\boxed{
\text{缺失模态类型}
+
\text{缺失位置}
+
\text{缺失长度}
}
\]
对预测性能的影响。
---
## 2.3 Q3:可解释性情感预测
在三模态信息完整条件下,不仅预测:
\[
(\hat c,\hat y)
\]
还需要解释:
1. 主要参考了哪种模态;
2. 不同模态的作用程度;
3. 哪些局部文本、音频和视觉片段最重要;
4. 关键证据对应原始视频的什么时间位置。
因此 Q3 的核心不是简单画 Attention,而是:
\[
\boxed{
\text{定位证据}
+
\text{量化贡献}
+
\text{验证解释}
+
\text{回溯原始素材}
}
\]
---
# 3. 数据集说明
---
## 3.1 附件 1:100 条原始视频
附件 1 包含从 CMU-MOSEI 中筛选的:
\[
100
\]
条英文视频样本。
共有:
\[
37
\]
个 `video_id` 子文件夹。
样本由:
```text
video_id + clip_id
```
共同唯一确定。
视频长度范围:
\[
2.648s \sim 34.567s
\]
配套文件:
```text
label-100.xlsx
```
包含字段:
| 字段 | 含义 |
|---|---|
| video_id | 原视频编号 |
| clip_id | 视频片段编号 |
| text | 英文转写文本 |
| label | 连续情感强度 |
| annotation | Negative / Neutral / Positive |
连续情感标签:
\[
y\in[-3,3]
\]
定义:
\[
[-3,0)\Rightarrow Negative
\]
\[
0\Rightarrow Neutral
\]
\[
(0,3]\Rightarrow Positive
\]
注意:
\[
\boxed{0\text{ 只属于 Neutral}}
\]
---
# 4. 附件 2:标准化多模态特征
附件 2 包含两套特征:
```text
aligned_50.pkl
unaligned_50.pkl
```
每套包含约:
\[
4850
\]
条有效样本,并按照:
```python
train
valid
test
```
划分。
---
## 4.1 Pickle 数据组织方式
整体结构为:
```text
data
├── train
│ ├── id
│ ├── raw_text
│ ├── text
│ ├── text_bert
│ ├── audio
│ ├── vision
│ ├── annotations
│ ├── classification_labels
│ ├── regression_labels
│ └── ...
├── valid
└── test
```
正确读取方式:
```python
data["train"]["audio"][j]
```
而不是:
```python
data["train"][j]["audio"]
```
---
# 5. aligned / unaligned 数据格式
## 5.1 aligned_50.pkl
\[
X^T\in\mathbb R^{N\times50\times768}
\]
\[
X^A\in\mathbb R^{N\times50\times74}
\]
\[
X^V\in\mathbb R^{N\times50\times35}
\]
三个模态已经被组织为:
\[
50
\]
个相互对应的位置。
可以理解为:
\[
T_i\leftrightarrow A_i\leftrightarrow V_i.
\]
---
## 5.2 unaligned_50.pkl
文本:
\[
X^T\in\mathbb R^{N\times50\times768}
\]
音频:
\[
X^A\in\mathbb R^{N\times500\times74}
\]
视觉:
\[
X^V\in\mathbb R^{N\times500\times35}
\]
同时提供:
```text
audio_lengths
vision_lengths
```
记录实际有效长度。
因此 unaligned 数据保留了:
\[
\boxed{
\text{文本较粗语义轴}
+
\text{音频细粒度时序}
+
\text{视觉细粒度时序}
}
\]
这为学习式 Cross-Attention 对齐提供了空间。
---
# 6. 附件 3:模态局部缺失测试集
附件 3:
- 无标签;
- 已处理多模态特征;
- 包含随机局部模态缺失;
- 缺失区间表现为连续位置全部置零。
用于 Q2 最终预测。
---
# 7. 附件 4:可解释性专项测试集
附件 4:
- 无标签;
- 三模态完整;
- 同时提供原始真实场景视频;
- 特征组织形式与附件 2 一致。
用于:
\[
\boxed{
\text{预测}
+
\text{模态贡献}
+
\text{关键证据定位}
}
\]
可以根据 Q1 保存的时间映射将关键位置重新定位到:
- 原始文本;
- 音频时间段;
- 视频关键帧。
---
# 8. 统一数学符号
一个样本记为:
\[
\mathcal X_i=
(X_i^T,X_i^A,X_i^V)
\]
其中:
- \(T\):Text
- \(A\):Audio
- \(V\):Vision
原始模态序列:
\[
X^m=
[x_1^m,\ldots,x_{L_m}^m],
\qquad
m\in\{T,A,V\}.
\]
---
## 8.1 基础符号
| 符号 | 含义 |
|---|---|
| \(x_j^m\) | 模态 \(m\) 原始第 \(j\) 个位置特征 |
| \(\tilde x_i^m\) | 软对齐到文本位置 \(i\) 后的特征 |
| \(h_i^m\) | 模态编码器生成的隐藏表示 |
| \(M_j^m\) | 原始位置是否有效 |
| \(A_{ij}^m\) | Cross-Attention 对齐权重 |
| \(R_i^m\) | 对齐后的局部模态可靠度 |
| \(\alpha_i^m\) | 融合时模态权重 |
| \(z_i\) | 第 \(i\) 个位置的多模态融合表示 |
| \(\beta_i\) | 第 \(i\) 个位置对整段情绪的时间重要性 |
| \(g\) | 整条视频最终表示 |
| \(\hat y\) | 情感强度预测 |
| \(\hat c\) | 情感极性预测 |
---
## Q1
@Q1.md
---
# 可能使用的开源库、工具与模型
| 类别 | 工具 / 模型 | 用途 | 优先级 |
|---|---|---|---|
| 深度学习 | PyTorch | 网络、训练、Cross-Attention | 必须 |
| GPU | CUDA / cuDNN | GPU 加速 | 必须 |
| 数值计算 | NumPy | 数值运算 | 必须 |
| 数据 | Pandas | CSV / XLSX / 表格 | 必须 |
| 科学计算 | SciPy | Pearson / 统计分析 | 必须 |
| ML | scikit-learn | F1、Accuracy、MAE 等 | 必须 |
| 配置 | PyYAML | 实验配置 | 推荐 |
| 配置 | Hydra / OmegaConf | 多实验管理 | 推荐 |
| 日志 | TensorBoard | loss / metric | 推荐 |
| 日志 | MLflow | 系统实验管理 | 可选 |
| 调参 | Optuna | 超参数搜索 | 推荐 |
| 视频 | FFmpeg | 拆音频、转码 | 必须 |
| 视频 | OpenCV | 抽帧与时间映射 | 必须 |
| 文本 | Hugging Face Transformers | BERT 等 | 高 |
| 文本 | BERT-base-uncased | 768 维表示 | 高 |
| ASR/时间戳 | WhisperX | 词级时间定位 | 推荐 |
| 强制对齐 | Montreal Forced Aligner | transcript-audio alignment | 推荐 |
| 音频 | openSMILE | 声学/韵律特征 | 高 |
| 音频 | librosa | 基础信号处理 | 推荐 |
| 音频 | torchaudio | PyTorch 音频 | 推荐 |
| 音频 PLM | wav2vec 2.0 | 深层语音表示 | 可选 |
| 音频 PLM | WavLM | 深层语音表示 | 可选 |
| 音频 PLM | HuBERT | 深层语音表示 | 可选 |
| 视觉 | OpenFace 2.0 | AU / gaze / pose | 高 |
| 视觉 | MediaPipe | landmark / blendshape | 可选 |
| 视觉 | torchvision | 图像处理 | 推荐 |
| 多模态 | CMU-MultimodalSDK | MOSEI 数据参考 | 高 |
| 多模态 | MultiBench | baseline / multimodal benchmark | 高 |
| 解释 | Captum | IG / Occlusion / Ablation | 高 |
| 解释 | SHAP | Shapley 解释 | 可选 |
| 可视化 | Matplotlib | 曲线 / 热图 | 必须 |
| 可视化 | Plotly | 交互可视化 | 可选 |
| 统计 | statsmodels | 回归 / 显著性分析 | 推荐 |
---
# 模型候选汇总
| 模块 | 候选 |
|---|---|
| Text Encoder | BERT |
| Audio Feature | openSMILE |
| Audio Encoder | BiGRU / Transformer |
| Audio PLM | Wav2Vec2 / WavLM |
| Vision Feature | OpenFace |
| Vision Encoder | BiGRU / Transformer |
| Alignment | Hard Alignment |
| Alignment | Fixed Window |
| Alignment | Cross-Attention |
| Fusion | Concatenation |
| Fusion | TFN |
| Fusion | LMF |
| Fusion | MAG |
| Fusion | MulT |
| Representation | MISA |
| Missing Model | MMIN |
| Missing Model | CMAD |
| Missing Model | P-RMF |
| Experts | EMOE |
| Diffusion | HyperEF |
| Factorization | FUSE-Net |
| Distribution Alignment | CaReFlow |
| Robust Representation | CmIR |
| Attribution | Integrated Gradients |
| Attribution | Occlusion |
| Attribution | Feature Ablation |
| Attribution | SHAP |
---
# 参考文献
> 下表优先使用论文官方会议/期刊入口。
> README 中记录 DOI、ACL Anthology ID、arXiv ID 或官方论文标题,避免引用二手博客作为正式论文来源。
| # | 文献 | 会议/期刊 | 年份 | 与本项目关系 | 官方标识 |
|---|---|---|---:|---|---|
| 1 | Zhang et al., *Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects* | Expert Systems with Applications | 2024 | 多模态情感综述 | DOI: 10.1016/j.eswa.2023.121692 |
| 2 | Bagher Zadeh et al., *Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph* | ACL | 2018 | CMU-MOSEI | ACL: P18-1208 |
| 3 | Vaswani et al., *Attention Is All You Need* | NeurIPS | 2017 | Transformer / Attention | NeurIPS 2017 |
| 4 | Devlin et al., *BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding* | NAACL | 2019 | 文本特征 | ACL: N19-1423 |
| 5 | Eyben et al., *openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor* | ACM Multimedia | 2010 | 音频特征 | ACM MM 2010 |
| 6 | Baltrušaitis et al., *OpenFace 2.0: Facial Behavior Analysis Toolkit* | FG | 2018 | 视觉特征 | DOI: 10.1109/FG.2018.00019 |
| 7 | Bain et al., *WhisperX: Time-Accurate Speech Transcription of Long-Form Audio* | arXiv | 2023 | 词级时间戳 | arXiv:2303.00747 |
| 8 | Tsai et al., *Multimodal Transformer for Unaligned Multimodal Language Sequences* | ACL | 2019 | unaligned + Cross-Attention | ACL: P19-1656 |
| 9 | Rahman et al., *Integrating Multimodal Information in Large Pretrained Transformers* | ACL | 2020 | MAG-BERT / 文本中心融合 | DOI:10.18653/v1/2020.acl-main.214 |
| 10 | Zadeh et al., *Tensor Fusion Network for Multimodal Sentiment Analysis* | EMNLP | 2017 | TFN baseline | ACL: D17-1115 |
| 11 | Liu et al., *Efficient Low-rank Multimodal Fusion With Modality-Specific Factors* | ACL | 2018 | LMF baseline | ACL: P18-1209 |
| 12 | Hazarika et al., *MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis* | ACM MM | 2020 | shared / private representation | DOI:10.1145/3394171.3413678 |
| 13 | Neverova et al., *ModDrop: Adaptive Multi-Modal Gesture Recognition* | TPAMI | 2016 | 模态随机丢失 | DOI:10.1109/TPAMI.2015.2461544 |
| 14 | Pham et al., *Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities* | AAAI | 2019 | 跨模态重构 | DOI:10.1609/aaai.v33i01.33016892 |
| 15 | Zhao et al., *Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities* | ACL-IJCNLP | 2021 | MMIN | ACL:2021.acl-long.203 |
| 16 | 王楠、王淇、欧阳丹彤,《基于知识蒸馏与动态调整机制的多模态情感分析模型》 | 计算机学报 | 2025 | AUMDF / 动态权重 / 缺失 | 48(8):1923–1942 |
| 17 | Fang et al., *EMOE: Modality-Specific Enhanced Dynamic Emotion Experts* | CVPR | 2025 | Mixture of Experts | CVPR 2025 |
| 18 | Zhuang et al., *CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities* | ICCV | 2025 | 缺失模态蒸馏 | ICCV 2025 |
| 19 | Zhu et al., *Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data* | ACL | 2025 | P-RMF / 不确定性 | DOI:10.18653/v1/2025.acl-long.1075 |
| 20 | Qiu et al., *Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition* | CVPR | 2026 | Diffusion / uncertainty | CVPR 2026 |
| 21 | Yang & Li, *Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis* | CVPR | 2026 | shared/specific/noise + reconstruction | CVPR 2026 |
| 22 | Mai & Han, *Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective* | ACL | 2026 | CmIR / causal invariance | DOI:10.18653/v1/2026.acl-long.2119 |
| 23 | Mai & Han, *CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion* | CVPR | 2026 | 模态分布对齐 | arXiv:2602.19140 |
| 24 | Wan et al., *Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation* | ACL | 2026 | 关键证据定位 | DOI:10.18653/v1/2026.acl-long.2012 |
| 25 | Sundararajan et al., *Axiomatic Attribution for Deep Networks* | ICML | 2017 | Integrated Gradients | PMLR 70 |
| 26 | Jain & Wallace, *Attention is not Explanation* | NAACL | 2019 | Attention 解释局限 | DOI:10.18653/v1/N19-1357 |
| 27 | Lundberg & Lee, *A Unified Approach to Interpreting Model Predictions* | NeurIPS | 2017 | SHAP | NeurIPS 2017 |
| 28 | Adebayo et al., *Sanity Checks for Saliency Maps* | NeurIPS | 2018 | 解释可信性验证 | NeurIPS 2018 |
| 29 | Zeiler & Fergus, *Visualizing and Understanding Convolutional Networks* | ECCV | 2014 | Occlusion 思想 | DOI:10.1007/978-3-319-10590-1_53 |
| 30 | Wachter et al., *Counterfactual Explanations without Opening the Black Box* | arXiv / HILDA | 2017 | 反事实解释 | arXiv:1711.00399 |
---
# 最终原则
本项目当前不采用:
> “先选一个看起来高级的模型,然后证明它最好。”
而采用:
\[
\boxed{
\text{提出多个机制假设}
\rightarrow
\text{设计公平实验}
\rightarrow
\text{观察数据}
\rightarrow
\text{解释规律}
\rightarrow
\text{确定最终模型}
}
\]
重点不只是得到最高指标,还要回答:
1. 为什么某种对齐方式更好?
2. 哪种模态最怕局部缺失?
3. 缺失的位置是否比缺失比例更重要?
4. Cross-Attention 能否帮助判断缺失信息的重要程度?
5. 动态可靠度是否真的提高鲁棒性?
6. 重构和“降低信任”哪种策略更有效?
7. Attention 权重是否真的对应模型决策依据?
8. 哪种解释方法最能经受反事实验证?
最终目标是形成:
\[
\boxed{
\text{有性能}
+
\text{有机制}
+
\text{有数学分析}
+
\text{有可解释性}
}
\]
的完整多模态情感预测建模方案。