提交其余项目实验变更

This commit is contained in:
2026-09-25 10:41:58 +08:00
parent 83ec3d1a83
commit 95bd34599b
119 changed files with 5877 additions and 1709 deletions
+67 -29
View File
@@ -1,39 +1,77 @@
# Q2/Q3 algorithm selection built on Q1 alignment
# Q2:缺失模态下的多模态情感识别
## Decision about reusing Q1
本目录包含 Q2 的训练代码、环境配置和后续实验约定。当前只维护两种模型:
The transferable part of Q1 is its explicit time correspondence and observation mask: features from different modalities share ordered positions, missing values are accompanied by masks, and a position can be traced to source time. That interface is useful for both Q2 local-gap handling and Q3 evidence localization.
1. **EarlyConcat + BiGRU**:将文本、音频、视觉特征和观测掩码拼接后,用双向 GRU 建模有序序列,作为简洁基线。
2. **MoFE-7 + MLP Router**:根据每个位置可用的模态,在七种模态子集专家之间路由,再用共享 BiGRU 建模序列。
The exact Q1 B1 extraction cannot be rerun over the 4,850 Attachment 2 training examples. Attachment 2 supplies precomputed aligned and unaligned feature tensors, but not the source audio/video or CTC word-time posteriors for the full training set. Its `aligned_50.pkl` also has 50 wordpiece positions and no Q1 `time_bounds_s`; those positions must not be described as the 50 equal-duration physical-time bins exported by `final/Q1`.
旧的模型比较报告、指标表和图表已清理。不要从此 README 推断模型优劣;后续结果应放进独立实验目录并在新报告中解释。
Accordingly, the Q2 experiment uses the official aligned feature set as its shared wordpiece axis, and compares it with a fixed equal-window pooling control made from the official unaligned audio/vision sequences. This is a downstream alignment-utility check, not a claim that B1 was recomputed on Attachment 2. The official train/validation split is retained; test labels are not used.
## 数据与时序表示
## Q2 candidates
训练入口读取项目根目录下 `E题数据/附件2-数据集特征文件/aligned_50.pkl` 的官方训练集和验证集。每个样本包含 50 个有序位置,文本、音频、视觉维度分别为 768、74、35,并带有显式观测掩码。这里的 50 个位置是附件 2 提供的词片位置,**不是 50 个等长物理时间箱**。
All candidates use identical training examples, train-only median/MAD scaling, joint polarity/intensity objectives, and 15 validation corruptions (three contiguous missing rates by five modality patterns).
Q1 的对齐方法为 Q2 提供了有序的跨模态输入组织方式;Q2 在此基础上处理连续块缺失,不重新提取或改写 Q1 特征。根目录 `math/` 和 `final/Q1/` 中的内容只作只读参考。
| Candidate | Fusion rule | What it tests |
| --- | --- | --- |
| `concat` | Project each modality, concatenate features and availability flags, then run a bidirectional GRU | Strong, simple early-fusion baseline |
| `gate` | Learn per-slot modality weights, mask unavailable modalities, then run a bidirectional GRU | Whether explicit reliability-aware fusion handles local gaps |
| `crossattn` | Apply masked cross-modal attention over the 50 shared slots, then temporal pooling | Whether contextual cross-modal exchange improves robustness |
标准化参数只从训练集拟合。当前保留的两组模型权重及共享标准化参数位于:
The report keeps Macro-F1, MAE, and Pearson separate. The default selection is Macro-F1-first across local corruption conditions; MAE and Pearson remain explicit tradeoffs, not terms in a constructed total score. The selected architecture is also trained on fixed-window-resampled features as an alignment control. A separate validation control shifts audio and vision by 1–10 positions to measure sensitivity to cross-modal timing.
## Q3 explanation selection
The selected Q2 model is frozen. Integrated Gradients and five-slot grouped occlusion are compared on held-out Attachment 2 validation clips using deletion comprehensiveness, sufficiency, and local rank stability. Attachment 4 has original videos and transcripts, so B1's CTC hard word-time procedure can be applied to those 20 clips to map high-importance wordpiece positions back to seconds. The saved Attachment 4 pickle files do not include `time_bounds_s`; explanations therefore retain both the model slot and the CTC-derived word interval, with alignment quality recorded.
## Run
The project environment is managed by `uv` and installs the CUDA 13.0 PyTorch build:
```bash
cd deep_learning/Q2
uv sync
uv run python -m q2.train_compare
cd ../Q3
uv run --project ../Q2 python -m q3.explain_selection
```text
outputs/mofe_7experts/
├── aligned_robust_stats.npz
└── models/
├── baselines/concat/seed_{42,3407,2026}/model_best.pt
└── B5_mofe_mlp/seed_{42,3407,2026}/model_best.pt
```
The main outputs are written to `outputs/algorithm_selection/`; plots, CSV metrics, run metadata, and checkpoints stay under this directory. The source data, `math`, and `final/Q1` are read-only inputs.
## 环境
本项目使用 `uv` 管理 Python 环境。进入本目录后同步锁定依赖:
```bash
uv sync --locked
```
## 运行
先用单个种子执行快速检查。每轮运行使用新的目录,避免覆盖保留的参照权重:
```bash
uv run python -m q2.train_mofe \
--phase smoke --seeds 42 \
--output-dir outputs/followups/F00_smoke
```
完整训练和验证示例:
```bash
uv run python -m q2.train_mofe \
--phase full --seeds 42 3407 2026 \
--output-dir outputs/followups/F01_local_repair
```
完整运行会训练两种模型,并在 clean、Text、Audio、Vision、Audio+Vision、All-modal 条件下评估 10%、20%、30% 连续块缺失;结果、检查点和诊断图写入指定目录。默认输出目录是 `outputs/mofe_7experts/`,日常新实验应显式设置 `--output-dir`,避免覆盖保留的权重和标准化参数。
## 文件索引
- `q2/data.py`:官方特征读取、观测掩码、训练集 robust scaling 和连续块缺失。
- `q2/models.py`:EarlyConcat + BiGRU。
- `q2/mofe.py`:MoFE-7 专家与 MLP 路由器。
- `q2/train_mofe.py`:两种保留模型的训练、验证、统计和可视化入口。
- `q2/task_preference.py`:复用现有 MoFE 权重,比较七个 forced expert 的分类/回归偏好。
- [ALGORITHM.md](ALGORITHM.md):两个保留模型的算法说明和最新验证集重评结果。
- [EXPERIMENT_PROTOCOL.md](EXPERIMENT_PROTOCOL.md):后续实验的固定比较条件与记录要求。
- `outputs/followups/README.md`:新实验目录的命名和存放规则。
Q3 暂不在本目录中开展;待 Q2 后续选型完成后再统一规划。
## 当前任务偏好诊断
在训练 dual-router 前,先用现有 MoFE-7 检查了七个 forced expert 在分类 Macro-F1 与回归 MAE/Pearson 上的偏好。三个 seed、15 种缺失条件的排名大体一致,没有看到稳定的分类—回归 expert 分工;因此当前不启动 dual-router,仍以 single-router MoFE-7 为活动参照。详见[诊断报告](outputs/followups/D0_task_preference/task_preference_diagnostic.md)及逐条件数据。
可用以下命令复现该诊断:
```bash
uv run python -m q2.task_preference \
--seeds 42 3407 2026 \
--output-dir outputs/followups/D0_task_preference
```