55 lines
5.3 KiB
Markdown
55 lines
5.3 KiB
Markdown
# Q2 algorithm selection results
|
||
|
||
## Q1 alignment transfer decision
|
||
|
||
Q1 B1 aligns BERT word features and audio/vision observations with hard CTC word intervals, projects observed features onto a 0.1-second common grid, exports 50 equal-duration physical-time bins, and keeps observation masks. For Q2, the shared ordered axis and explicit masks transfer directly: a local gap stays a local gap after alignment and can be represented without inventing feature values.
|
||
|
||
The exact B1 extraction was not recomputed over Attachment 2. The official 4,850-row feature package contains precomputed aligned and unaligned tensors, but no full-set source audio/video or word-time posterior. Its `aligned_50.pkl` has 50 wordpiece positions and no per-slot `time_bounds_s`; those positions are not Q1's 50 equal-duration bins. This experiment therefore trains on the official aligned features and compares them with an equal-window audio/vision resampling control. The comparison tests the value of an aligned ordered representation for the downstream Q2 task; it does not claim to reproduce B1 on all 4,850 clips.
|
||
|
||
## Data and protocol
|
||
|
||
- Attachment 2 official split: 3,395 training clips and 728 validation clips. Their source-video ID sets do not overlap.
|
||
- Each official aligned example has 50 positions with Text 768-D, Audio 74-D, Vision 35-D features and modality observation masks.
|
||
- Attachment 2 test labels were not used.
|
||
- The three candidates shared train-only median/MAD normalization, the joint polarity/intensity objective, and training-time contiguous block masking.
|
||
- Validation corruption covered 10%, 20%, and 30% of 50 positions for Text, Audio, Vision, Audio+Vision, and all three modalities. This is a wordpiece-position proxy for a continuous time gap; full-set second-level timestamps are not supplied.
|
||
- Each candidate was run with seeds 42, 3407, and 2026. Reported `±` values are seed standard deviations over the fixed official validation set and deterministic corruption draws; they are not confidence intervals over new videos.
|
||
|
||
## Fusion comparison
|
||
|
||
| Model | Clean Accuracy | Clean Macro-F1 | Corrupt Accuracy, mean | Corrupt Macro-F1, mean | Worst condition Macro-F1 | Corrupt MAE | Corrupt Pearson |
|
||
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||
| Early concatenation + BiGRU | 0.626 ± 0.013 | 0.580 ± 0.018 | 0.623 ± 0.011 | **0.575 ± 0.017** | **0.516** | **0.640 ± 0.004** | 0.607 ± 0.004 |
|
||
| Reliability gate + BiGRU | 0.621 ± 0.011 | 0.570 ± 0.019 | 0.617 ± 0.011 | 0.565 ± 0.018 | 0.498 | 0.643 ± 0.008 | **0.610 ± 0.008** |
|
||
| Masked cross-modal attention | 0.604 ± 0.012 | 0.540 ± 0.049 | 0.601 ± 0.009 | 0.539 ± 0.047 | 0.462 | 0.649 ± 0.024 | 0.592 ± 0.016 |
|
||
|
||
Early concatenation has the best mean corrupted Macro-F1 and MAE. The gate has slightly higher Pearson, so the metrics do not collapse to one score. Cross-modal attention is lower and more variable at this sample size. It is not selected for the next Q2 stage.
|
||
|
||
For the selected concatenation model, the hardest tested case is 30% Text masking: Macro-F1 0.547 and MAE 0.665, compared with clean Macro-F1 0.580 and MAE 0.636. Audio-only or Vision-only masking has a smaller effect in these runs. This is evidence about this feature set and these simulated spans; it does not establish a universal modality ranking.
|
||
|
||
## Alignment utility control
|
||
|
||
The same concatenation model was trained either on the supplied aligned wordpiece features or on equal-window-resampled audio/vision features from the official unaligned tensors.
|
||
|
||
| Representation | Clean Macro-F1 | Corrupt Macro-F1 | Corrupt MAE | Corrupt Pearson |
|
||
| --- | ---: | ---: | ---: | ---: |
|
||
| Supplied word-aligned 50 positions | 0.580 ± 0.018 | **0.575 ± 0.017** | **0.640 ± 0.004** | **0.607 ± 0.004** |
|
||
| Equal-window resampled unaligned input | 0.501 ± 0.014 | 0.504 ± 0.016 | 0.665 ± 0.005 | 0.577 ± 0.010 |
|
||
|
||
On the selected model, shifting Audio and Vision by 1–10 positions changed aligned Macro-F1 from 0.580 to 0.557. That is a modest timing-sensitivity signal; it does not prove the model uses precise physical-time correspondence. Together with the fixed-window comparison, the result supports retaining the supplied aligned sequence for Q2.
|
||
|
||
## Selected Q2 direction
|
||
|
||
Continue with mask-aware early concatenation plus a bidirectional GRU, using local block masking during training. Keep the reliability gate as an ablation because its Pearson is slightly higher. Revisit cross-attention only if a later run has stronger evidence and enough data to control overfitting.
|
||
|
||
## Reproducible artifacts
|
||
|
||
- [Model and representation summary](outputs/algorithm_selection/summary.csv)
|
||
- [Metrics by missing type and rate](outputs/algorithm_selection/validation_metrics_by_condition.csv)
|
||
- [Aligned versus fixed-window and temporal-shift controls](outputs/algorithm_selection/alignment_transfer_ablation.csv)
|
||
- [Training/data audit and run manifest](outputs/algorithm_selection/data_audit.json), [run manifest](outputs/algorithm_selection/run_manifest.json)
|
||
- [Validation plot](outputs/algorithm_selection/missing_rate_comparison.png)
|
||
- [Selected seed-42 checkpoint](outputs/algorithm_selection/models/aligned/concat/model_best.pt)
|
||
|
||
The fitted checkpoint is for algorithm selection, not the final Attachment 3 submission model. The final model should be trained on train+validation after the architecture and thresholds are frozen.
|