Q2/Q3 algorithm selection built on Q1 alignment
Decision about reusing Q1
The transferable part of Q1 is its explicit time correspondence and observation mask: features from different modalities share ordered positions, missing values are accompanied by masks, and a position can be traced to source time. That interface is useful for both Q2 local-gap handling and Q3 evidence localization.
The exact Q1 B1 extraction cannot be rerun over the 4,850 Attachment 2 training examples. Attachment 2 supplies precomputed aligned and unaligned feature tensors, but not the source audio/video or CTC word-time posteriors for the full training set. Its aligned_50.pkl also has 50 wordpiece positions and no Q1 time_bounds_s; those positions must not be described as the 50 equal-duration physical-time bins exported by final/Q1.
Accordingly, the Q2 experiment uses the official aligned feature set as its shared wordpiece axis, and compares it with a fixed equal-window pooling control made from the official unaligned audio/vision sequences. This is a downstream alignment-utility check, not a claim that B1 was recomputed on Attachment 2. The official train/validation split is retained; test labels are not used.
Q2 candidates
All candidates use identical training examples, train-only median/MAD scaling, joint polarity/intensity objectives, and 15 validation corruptions (three contiguous missing rates by five modality patterns).
| Candidate | Fusion rule | What it tests |
|---|---|---|
concat |
Project each modality, concatenate features and availability flags, then run a bidirectional GRU | Strong, simple early-fusion baseline |
gate |
Learn per-slot modality weights, mask unavailable modalities, then run a bidirectional GRU | Whether explicit reliability-aware fusion handles local gaps |
crossattn |
Apply masked cross-modal attention over the 50 shared slots, then temporal pooling | Whether contextual cross-modal exchange improves robustness |
The report keeps Macro-F1, MAE, and Pearson separate. The default selection is Macro-F1-first across local corruption conditions; MAE and Pearson remain explicit tradeoffs, not terms in a constructed total score. The selected architecture is also trained on fixed-window-resampled features as an alignment control. A separate validation control shifts audio and vision by 1–10 positions to measure sensitivity to cross-modal timing.
Q3 explanation selection
The selected Q2 model is frozen. Integrated Gradients and five-slot grouped occlusion are compared on held-out Attachment 2 validation clips using deletion comprehensiveness, sufficiency, and local rank stability. Attachment 4 has original videos and transcripts, so B1's CTC hard word-time procedure can be applied to those 20 clips to map high-importance wordpiece positions back to seconds. The saved Attachment 4 pickle files do not include time_bounds_s; explanations therefore retain both the model slot and the CTC-derived word interval, with alignment quality recorded.
Run
The project environment is managed by uv and installs the CUDA 13.0 PyTorch build:
cd deep_learning/Q2
uv sync
uv run python -m q2.train_compare
cd ../Q3
uv run --project ../Q2 python -m q3.explain_selection
The main outputs are written to outputs/algorithm_selection/; plots, CSV metrics, run metadata, and checkpoints stay under this directory. The source data, math, and final/Q1 are read-only inputs.