5.3 KiB
Q2 algorithm selection results
Q1 alignment transfer decision
Q1 B1 aligns BERT word features and audio/vision observations with hard CTC word intervals, projects observed features onto a 0.1-second common grid, exports 50 equal-duration physical-time bins, and keeps observation masks. For Q2, the shared ordered axis and explicit masks transfer directly: a local gap stays a local gap after alignment and can be represented without inventing feature values.
The exact B1 extraction was not recomputed over Attachment 2. The official 4,850-row feature package contains precomputed aligned and unaligned tensors, but no full-set source audio/video or word-time posterior. Its aligned_50.pkl has 50 wordpiece positions and no per-slot time_bounds_s; those positions are not Q1's 50 equal-duration bins. This experiment therefore trains on the official aligned features and compares them with an equal-window audio/vision resampling control. The comparison tests the value of an aligned ordered representation for the downstream Q2 task; it does not claim to reproduce B1 on all 4,850 clips.
Data and protocol
- Attachment 2 official split: 3,395 training clips and 728 validation clips. Their source-video ID sets do not overlap.
- Each official aligned example has 50 positions with Text 768-D, Audio 74-D, Vision 35-D features and modality observation masks.
- Attachment 2 test labels were not used.
- The three candidates shared train-only median/MAD normalization, the joint polarity/intensity objective, and training-time contiguous block masking.
- Validation corruption covered 10%, 20%, and 30% of 50 positions for Text, Audio, Vision, Audio+Vision, and all three modalities. This is a wordpiece-position proxy for a continuous time gap; full-set second-level timestamps are not supplied.
- Each candidate was run with seeds 42, 3407, and 2026. Reported
±values are seed standard deviations over the fixed official validation set and deterministic corruption draws; they are not confidence intervals over new videos.
Fusion comparison
| Model | Clean Accuracy | Clean Macro-F1 | Corrupt Accuracy, mean | Corrupt Macro-F1, mean | Worst condition Macro-F1 | Corrupt MAE | Corrupt Pearson |
|---|---|---|---|---|---|---|---|
| Early concatenation + BiGRU | 0.626 ± 0.013 | 0.580 ± 0.018 | 0.623 ± 0.011 | 0.575 ± 0.017 | 0.516 | 0.640 ± 0.004 | 0.607 ± 0.004 |
| Reliability gate + BiGRU | 0.621 ± 0.011 | 0.570 ± 0.019 | 0.617 ± 0.011 | 0.565 ± 0.018 | 0.498 | 0.643 ± 0.008 | 0.610 ± 0.008 |
| Masked cross-modal attention | 0.604 ± 0.012 | 0.540 ± 0.049 | 0.601 ± 0.009 | 0.539 ± 0.047 | 0.462 | 0.649 ± 0.024 | 0.592 ± 0.016 |
Early concatenation has the best mean corrupted Macro-F1 and MAE. The gate has slightly higher Pearson, so the metrics do not collapse to one score. Cross-modal attention is lower and more variable at this sample size. It is not selected for the next Q2 stage.
For the selected concatenation model, the hardest tested case is 30% Text masking: Macro-F1 0.547 and MAE 0.665, compared with clean Macro-F1 0.580 and MAE 0.636. Audio-only or Vision-only masking has a smaller effect in these runs. This is evidence about this feature set and these simulated spans; it does not establish a universal modality ranking.
Alignment utility control
The same concatenation model was trained either on the supplied aligned wordpiece features or on equal-window-resampled audio/vision features from the official unaligned tensors.
| Representation | Clean Macro-F1 | Corrupt Macro-F1 | Corrupt MAE | Corrupt Pearson |
|---|---|---|---|---|
| Supplied word-aligned 50 positions | 0.580 ± 0.018 | 0.575 ± 0.017 | 0.640 ± 0.004 | 0.607 ± 0.004 |
| Equal-window resampled unaligned input | 0.501 ± 0.014 | 0.504 ± 0.016 | 0.665 ± 0.005 | 0.577 ± 0.010 |
On the selected model, shifting Audio and Vision by 1–10 positions changed aligned Macro-F1 from 0.580 to 0.557. That is a modest timing-sensitivity signal; it does not prove the model uses precise physical-time correspondence. Together with the fixed-window comparison, the result supports retaining the supplied aligned sequence for Q2.
Selected Q2 direction
Continue with mask-aware early concatenation plus a bidirectional GRU, using local block masking during training. Keep the reliability gate as an ablation because its Pearson is slightly higher. Revisit cross-attention only if a later run has stronger evidence and enough data to control overfitting.
Reproducible artifacts
- Model and representation summary
- Metrics by missing type and rate
- Aligned versus fixed-window and temporal-shift controls
- Training/data audit and run manifest, run manifest
- Validation plot
- Selected seed-42 checkpoint
The fitted checkpoint is for algorithm selection, not the final Attachment 3 submission model. The final model should be trained on train+validation after the architecture and thresholds are frozen.