Files
modeling_zhaocui/deep_learning/Q2

Q2/Q3 algorithm selection built on Q1 alignment

Decision about reusing Q1

The transferable part of Q1 is its explicit time correspondence and observation mask: features from different modalities share ordered positions, missing values are accompanied by masks, and a position can be traced to source time. That interface is useful for both Q2 local-gap handling and Q3 evidence localization.

The exact Q1 B1 extraction cannot be rerun over the 4,850 Attachment 2 training examples. Attachment 2 supplies precomputed aligned and unaligned feature tensors, but not the source audio/video or CTC word-time posteriors for the full training set. Its aligned_50.pkl also has 50 wordpiece positions and no Q1 time_bounds_s; those positions must not be described as the 50 equal-duration physical-time bins exported by final/Q1.

Accordingly, the Q2 experiment uses the official aligned feature set as its shared wordpiece axis, and compares it with a fixed equal-window pooling control made from the official unaligned audio/vision sequences. This is a downstream alignment-utility check, not a claim that B1 was recomputed on Attachment 2. The official train/validation split is retained; test labels are not used.

Q2 candidates

All candidates use identical training examples, train-only median/MAD scaling, joint polarity/intensity objectives, and 15 validation corruptions (three contiguous missing rates by five modality patterns).

Candidate Fusion rule What it tests
concat Project each modality, concatenate features and availability flags, then run a bidirectional GRU Strong, simple early-fusion baseline
gate Learn per-slot modality weights, mask unavailable modalities, then run a bidirectional GRU Whether explicit reliability-aware fusion handles local gaps
crossattn Apply masked cross-modal attention over the 50 shared slots, then temporal pooling Whether contextual cross-modal exchange improves robustness

The report keeps Macro-F1, MAE, and Pearson separate. The default selection is Macro-F1-first across local corruption conditions; MAE and Pearson remain explicit tradeoffs, not terms in a constructed total score. The selected architecture is also trained on fixed-window-resampled features as an alignment control. A separate validation control shifts audio and vision by 1–10 positions to measure sensitivity to cross-modal timing.

Q3 explanation selection

The selected Q2 model is frozen. Integrated Gradients and five-slot grouped occlusion are compared on held-out Attachment 2 validation clips using deletion comprehensiveness, sufficiency, and local rank stability. Attachment 4 has original videos and transcripts, so B1's CTC hard word-time procedure can be applied to those 20 clips to map high-importance wordpiece positions back to seconds. The saved Attachment 4 pickle files do not include time_bounds_s; explanations therefore retain both the model slot and the CTC-derived word interval, with alignment quality recorded.

Run

The project environment is managed by uv and installs the CUDA 13.0 PyTorch build:

cd deep_learning/Q2
uv sync
uv run python -m q2.train_compare
cd ../Q3
uv run --project ../Q2 python -m q3.explain_selection

The main outputs are written to outputs/algorithm_selection/; plots, CSV metrics, run metadata, and checkpoints stay under this directory. The source data, math, and final/Q1 are read-only inputs.