Files
modeling_zhaocui/deep_learning/Q2/README.md
T

40 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Q2/Q3 algorithm selection built on Q1 alignment
## Decision about reusing Q1
The transferable part of Q1 is its explicit time correspondence and observation mask: features from different modalities share ordered positions, missing values are accompanied by masks, and a position can be traced to source time. That interface is useful for both Q2 local-gap handling and Q3 evidence localization.
The exact Q1 B1 extraction cannot be rerun over the 4,850 Attachment 2 training examples. Attachment 2 supplies precomputed aligned and unaligned feature tensors, but not the source audio/video or CTC word-time posteriors for the full training set. Its `aligned_50.pkl` also has 50 wordpiece positions and no Q1 `time_bounds_s`; those positions must not be described as the 50 equal-duration physical-time bins exported by `final/Q1`.
Accordingly, the Q2 experiment uses the official aligned feature set as its shared wordpiece axis, and compares it with a fixed equal-window pooling control made from the official unaligned audio/vision sequences. This is a downstream alignment-utility check, not a claim that B1 was recomputed on Attachment 2. The official train/validation split is retained; test labels are not used.
## Q2 candidates
All candidates use identical training examples, train-only median/MAD scaling, joint polarity/intensity objectives, and 15 validation corruptions (three contiguous missing rates by five modality patterns).
| Candidate | Fusion rule | What it tests |
| --- | --- | --- |
| `concat` | Project each modality, concatenate features and availability flags, then run a bidirectional GRU | Strong, simple early-fusion baseline |
| `gate` | Learn per-slot modality weights, mask unavailable modalities, then run a bidirectional GRU | Whether explicit reliability-aware fusion handles local gaps |
| `crossattn` | Apply masked cross-modal attention over the 50 shared slots, then temporal pooling | Whether contextual cross-modal exchange improves robustness |
The report keeps Macro-F1, MAE, and Pearson separate. The default selection is Macro-F1-first across local corruption conditions; MAE and Pearson remain explicit tradeoffs, not terms in a constructed total score. The selected architecture is also trained on fixed-window-resampled features as an alignment control. A separate validation control shifts audio and vision by 1–10 positions to measure sensitivity to cross-modal timing.
## Q3 explanation selection
The selected Q2 model is frozen. Integrated Gradients and five-slot grouped occlusion are compared on held-out Attachment 2 validation clips using deletion comprehensiveness, sufficiency, and local rank stability. Attachment 4 has original videos and transcripts, so B1's CTC hard word-time procedure can be applied to those 20 clips to map high-importance wordpiece positions back to seconds. The saved Attachment 4 pickle files do not include `time_bounds_s`; explanations therefore retain both the model slot and the CTC-derived word interval, with alignment quality recorded.
## Run
The project environment is managed by `uv` and installs the CUDA 13.0 PyTorch build:
```bash
cd deep_learning/Q2
uv sync
uv run python -m q2.train_compare
cd ../Q3
uv run --project ../Q2 python -m q3.explain_selection
```
The main outputs are written to `outputs/algorithm_selection/`; plots, CSV metrics, run metadata, and checkpoints stay under this directory. The source data, `math`, and `final/Q1` are read-only inputs.