# Q2/Q3 algorithm selection built on Q1 alignment ## Decision about reusing Q1 The transferable part of Q1 is its explicit time correspondence and observation mask: features from different modalities share ordered positions, missing values are accompanied by masks, and a position can be traced to source time. That interface is useful for both Q2 local-gap handling and Q3 evidence localization. The exact Q1 B1 extraction cannot be rerun over the 4,850 Attachment 2 training examples. Attachment 2 supplies precomputed aligned and unaligned feature tensors, but not the source audio/video or CTC word-time posteriors for the full training set. Its `aligned_50.pkl` also has 50 wordpiece positions and no Q1 `time_bounds_s`; those positions must not be described as the 50 equal-duration physical-time bins exported by `final/Q1`. Accordingly, the Q2 experiment uses the official aligned feature set as its shared wordpiece axis, and compares it with a fixed equal-window pooling control made from the official unaligned audio/vision sequences. This is a downstream alignment-utility check, not a claim that B1 was recomputed on Attachment 2. The official train/validation split is retained; test labels are not used. ## Q2 candidates All candidates use identical training examples, train-only median/MAD scaling, joint polarity/intensity objectives, and 15 validation corruptions (three contiguous missing rates by five modality patterns). | Candidate | Fusion rule | What it tests | | --- | --- | --- | | `concat` | Project each modality, concatenate features and availability flags, then run a bidirectional GRU | Strong, simple early-fusion baseline | | `gate` | Learn per-slot modality weights, mask unavailable modalities, then run a bidirectional GRU | Whether explicit reliability-aware fusion handles local gaps | | `crossattn` | Apply masked cross-modal attention over the 50 shared slots, then temporal pooling | Whether contextual cross-modal exchange improves robustness | The report keeps Macro-F1, MAE, and Pearson separate. The default selection is Macro-F1-first across local corruption conditions; MAE and Pearson remain explicit tradeoffs, not terms in a constructed total score. The selected architecture is also trained on fixed-window-resampled features as an alignment control. A separate validation control shifts audio and vision by 1–10 positions to measure sensitivity to cross-modal timing. ## Q3 explanation selection The selected Q2 model is frozen. Integrated Gradients and five-slot grouped occlusion are compared on held-out Attachment 2 validation clips using deletion comprehensiveness, sufficiency, and local rank stability. Attachment 4 has original videos and transcripts, so B1's CTC hard word-time procedure can be applied to those 20 clips to map high-importance wordpiece positions back to seconds. The saved Attachment 4 pickle files do not include `time_bounds_s`; explanations therefore retain both the model slot and the CTC-derived word interval, with alignment quality recorded. ## Run The project environment is managed by `uv` and installs the CUDA 13.0 PyTorch build: ```bash cd deep_learning/Q2 uv sync uv run python -m q2.train_compare cd ../Q3 uv run --project ../Q2 python -m q3.explain_selection ``` The main outputs are written to `outputs/algorithm_selection/`; plots, CSV metrics, run metadata, and checkpoints stay under this directory. The source data, `math`, and `final/Q1` are read-only inputs.